Mirinex

Credits

Mirinex is built on open-source work. Thank you to everyone below.

Separation model

Voice and music separation uses the MelBand RoFormer vocal model by Kimberley Jensen, released under the MIT licence. Source and licence.

Speech cleanup uses ClearerVoice and MossFormer2 SE 48K by Alibaba Inc., under the Apache License 2.0. We preserve timing with overlapping inference windows.

Four-stem separation combines the vocal model above with SCNet XL IHF weights released under MIT by Roman Solovyev (ZFTurbo). Its SCNet implementation is also MIT; full notice. We apply overlapping windows, residual reconstruction and one shared gain for the exported tracks.

In your browser

Sample song

The vocal remover's sample song was made for Mirinex with Artlist's AI music. Its vocals and music were separated by Mirinex.

Captions

Speech recognition uses Whisper large-v3-turbo by OpenAI (MIT licence), run through WhisperX (BSD-2-Clause), faster-whisper (MIT) and CTranslate2 (MIT). Word timing uses wav2vec 2.0 alignment models: torchaudio's English model (MIT); Jonatas Grosman's XLSR-53 models for Spanish, French, German, Italian, Portuguese, Japanese, Chinese, Dutch, Arabic, Russian, Polish, Hungarian, Finnish, Persian and Greek (Apache 2.0); and community models for other languages under Apache 2.0 or similar open licences, including mpoyraz's Turkish model (CC BY 4.0) and KBLab's Swedish model (CC0). Speaker labels use pyannote's speaker-diarization-community-1 pipeline by pyannoteAI, under CC BY 4.0.

Images

On our server