Credits
Mirinex is built on open-source work. Thank you to everyone below.
Separation model
Voice and music separation uses the MelBand RoFormer vocal model by Kimberley Jensen, released under the MIT licence. Source and licence.
Speech cleanup uses ClearerVoice and MossFormer2 SE 48K by Alibaba Inc., under the Apache License 2.0. We preserve timing with overlapping inference windows.
Four-stem separation combines the vocal model above with SCNet XL IHF weights released under MIT by Roman Solovyev (ZFTurbo). Its SCNet implementation is also MIT; full notice. We apply overlapping windows, residual reconstruction and one shared gain for the exported tracks.
In your browser
- Mediabunny by Vanilagy reads and writes audio and video files in the page. Mozilla Public License 2.0; our copy is unmodified, and its licence is at /prototype/node_modules/mediabunny/LICENSE.
- MP3 encoding uses LAME 3.100 under the GNU Lesser General Public License, through the unmodified Mediabunny MP3 encoder 1.61.0 (MPL-2.0). LAME licence · LAME source · Encoder source/build instructions. The encoder is supplied as a separate module and may be replaced or modified under these licences.
- Tempo analysis uses the BeatRoot implementation in music-tempo 1.0.3, Copyright (c) 2017 killercrush. MIT licence. Key estimation uses published Krumhansl–Kessler pitch-class profiles.
- Loudness measurement adapts filter and true-peak interpolation code from libebur128 1.2.6, Copyright (c) 2011 Jan Kokemüller, under the MIT licence. Our browser implementation.
- Pitch and tempo use Signalsmith Stretch, official browser release 1.3.2, Copyright (c) 2022 Geraint Luff / Signalsmith Audio Ltd., under the MIT licence. We expose its existing WASM factory for offline workers; the algorithm is unchanged. Original release · Adapted module.
- Inter by Rasmus Andersson, SIL Open Font License 1.1. Licence.
Sample song
The vocal remover's sample song was made for Mirinex with Artlist's AI music. Its vocals and music were separated by Mirinex.
Captions
Speech recognition uses Whisper large-v3-turbo by OpenAI (MIT licence), run through WhisperX (BSD-2-Clause), faster-whisper (MIT) and CTranslate2 (MIT). Word timing uses wav2vec 2.0 alignment models: torchaudio's English model (MIT); Jonatas Grosman's XLSR-53 models for Spanish, French, German, Italian, Portuguese, Japanese, Chinese, Dutch, Arabic, Russian, Polish, Hungarian, Finnish, Persian and Greek (Apache 2.0); and community models for other languages under Apache 2.0 or similar open licences, including mpoyraz's Turkish model (CC BY 4.0) and KBLab's Swedish model (CC0). Speaker labels use pyannote's speaker-diarization-community-1 pipeline by pyannoteAI, under CC BY 4.0.
Images
- The background remover's animated bust of Michelangelo's David is made from David (Michelangelo), a 3D scan by Scan the World, licensed under CC BY-SA 4.0. We cropped it to a bust, removed the hand, closed its base and simplified it; our adapted model is shared under the same licence.
On our server
- python-audio-separator (MIT), PyTorch (BSD-3-Clause) and FastAPI (MIT).