Mirinex

Technology

The engine behind
every tool.

Mirinex runs state-of-the-art separation, speech, captioning and vision models on dedicated NVIDIA GPUs, with our own processing around them for clean, full-length, studio-quality results.

11.0 dBVocal SDR on the public MVSEP Multisong benchmark
17.3 dBInstrumental SDR on the same benchmark
99Caption languages, aligned word by word
24-bitLossless WAV and FLAC export, up to 48 kHz

SDR (signal-to-distortion ratio) measures how close a separated track is to the real one; higher is better. Scores are the published results for our vocal model on MVSEP.

The models

Vocals and music

Mel-Band RoFormer

ByteDance research architecture, weights by Kimberley Jensen (MIT).

Every song and video is separated at full length, up to 20 minutes, with overlapping windows at full resolution and 24-bit output. The same model family leads the public separation leaderboards.

Four stems

SCNet XL

Weights and implementation by Roman Solovyev (ZFTurbo), MIT.

Combined with the vocal model to give vocals, drums, bass and everything else, with residual reconstruction and one shared gain, so the stems add back up to the original mix.

Speech cleanup

MossFormer2 SE 48K

ClearerVoice by Alibaba, Apache 2.0.

Full-band 48 kHz enhancement for hiss, hum and room noise, run in overlapping windows so every word keeps its exact timing.

Captions

Whisper large-v3-turbo + WhisperX

OpenAI (MIT), WhisperX (BSD), wav2vec 2.0 aligners.

Voice activity detection, batched recognition in 99 languages, then forced alignment of every word to the audio, so word-by-word captions land exactly on the beat of speech.

Background removal

BiRefNet

Bilateral reference network for high-resolution segmentation, MIT.

Hair, fur and fine edges cut cleanly at up to 4 megapixels, with white backgrounds and square framing for product listings.

In your browser

WebCodecs + WebAssembly

Mediabunny, Signalsmith Stretch, EBU R128 loudness.

Decoding, video export, pitch and tempo, loudness and key analysis run on your own device, so editing tools respond instantly and those files never leave it.

From your file to your result

  1. Read in your browser. Audio and video are decoded on your device, and only the sound a tool needs is sent.
  2. Processed on dedicated GPUs. Each file runs at full length on NVIDIA A10G GPUs, with automatic retries if anything fails.
  3. Heard before you download. Results play in the page, where you can compare, mix and adjust them first.
  4. Exported at full quality. Lossless audio, the original video with the new sound, captioned video, or PNG.

Private by design

Deleted automatically

Uploaded files are deleted automatically, and results expire shortly after they’re ready. Pro keeps a 30-day library of your own results.

Never used for training

Your files are processed to give you your result, and never used to train models.

Encrypted in transit

Every upload and download travels over HTTPS, and payments are handled by Stripe.

On your device where possible

Editing, analysis and export run in your browser, so those files never leave your device.

Hear it on your own file.

Choose a file