Vocals and music
Mel-Band RoFormer
ByteDance research architecture, weights by Kimberley Jensen (MIT).
Every song and video is separated at full length, up to 20 minutes, with overlapping windows at full resolution and 24-bit output. The same model family leads the public separation leaderboards.
Four stems
SCNet XL
Weights and implementation by Roman Solovyev (ZFTurbo), MIT.
Combined with the vocal model to give vocals, drums, bass and everything else, with residual reconstruction and one shared gain, so the stems add back up to the original mix.
Speech cleanup
MossFormer2 SE 48K
ClearerVoice by Alibaba, Apache 2.0.
Full-band 48 kHz enhancement for hiss, hum and room noise, run in overlapping windows so every word keeps its exact timing.
Captions
Whisper large-v3-turbo + WhisperX
OpenAI (MIT), WhisperX (BSD), wav2vec 2.0 aligners.
Voice activity detection, batched recognition in 99 languages, then forced alignment of every word to the audio, so word-by-word captions land exactly on the beat of speech.
Background removal
BiRefNet
Bilateral reference network for high-resolution segmentation, MIT.
Hair, fur and fine edges cut cleanly at up to 4 megapixels, with white backgrounds and square framing for product listings.
In your browser
WebCodecs + WebAssembly
Mediabunny, Signalsmith Stretch, EBU R128 loudness.
Decoding, video export, pitch and tempo, loudness and key analysis run on your own device, so editing tools respond instantly and those files never leave it.