--- license: mit pipeline_tag: audio-to-audio tags: - audio-separation - vocal-remover - stem-separation - onnx - onnxruntime --- # AudioSeparatorONNX Collection of popular audio separation and vocal removal models converted to **ONNX** format, optimized for high-performance CPU/GPU inference using `onnxruntime`. ## 📌 Overview This repository provides pre-converted ONNX versions of popular music source separation (MSS) and vocal removal models (such as UVR5 / Ultimate Vocal Remover models, MDX-Net, VR Architecture, Demucs, etc.). Using ONNX models allows for: - **Cross-platform compatibility** (Python, C++, C#, Rust, JS, etc.). - **Faster execution** with `onnxruntime` using CPU, CUDA, DirectML, CoreML, or TensorRT. - **Lightweight deployments** without requiring heavy frameworks like PyTorch or TensorFlow in production. --- ## 💻 Quickstart ## 📦 Embedded Metadata (`sep_meta`) Every `.onnx` file in this repository is **self-contained** — no separate sidecar `.json` file is needed. All inference parameters are embedded directly inside the ONNX model as a `metadata_props` entry with key `sep_meta`. ### Reading metadata with Python (`onnxruntime`) ```python import json import onnxruntime as ort sess = ort.InferenceSession("Reverb_HQ_By_FoxJoy.onnx", providers=["CPUExecutionProvider"]) meta_map = sess.get_modelmeta().custom_metadata_map # dict[str, str] sep_meta = json.loads(meta_map["sep_meta"]) print(sep_meta["arch"]) # "MDX" | "VR" | "ROFORMER" print(sep_meta["primary_stem"]) # "No Reverb" print(sep_meta["secondary_stem"]) # "Reverb" ``` ### Reading metadata with Python (`onnx` library, no ORT session) ```python import json import onnx model = onnx.load("UVR-MDX-NET-Inst_HQ_1.onnx") meta_map = {p.key: p.value for p in model.metadata_props} sep_meta = json.loads(meta_map["sep_meta"]) ``` ### Reading metadata with C (fast byte-scan, no ORT dependency) The `sep_meta` key is stored near the **end** of the ONNX protobuf binary, after all weight initializers. You can extract it with a simple byte scan — no need to load the entire model into memory: ```c #include #include #include /* Returns a heap-allocated JSON string, or NULL on failure. Caller must free() the result. */ char* read_onnx_sep_meta(const char* onnx_path) { FILE* f = fopen(onnx_path, "rb"); if (!f) return NULL; /* Read the last 64 KB — sep_meta is always at the end */ const size_t SCAN_SIZE = 64 * 1024; fseek(f, 0, SEEK_END); long file_size = ftell(f); size_t read_size = (file_size < (long)SCAN_SIZE) ? (size_t)file_size : SCAN_SIZE; long offset = file_size - (long)read_size; fseek(f, offset, SEEK_SET); char* buf = (char*)malloc(read_size + 1); if (!buf) { fclose(f); return NULL; } size_t n = fread(buf, 1, read_size, f); fclose(f); buf[n] = '\0'; /* Find the marker string "sep_meta" */ const char* marker = "sep_meta"; char* pos = buf; char* found = NULL; while ((pos = (char*)memchr(pos, 's', buf + n - pos)) != NULL) { if ((size_t)(buf + n - pos) >= strlen(marker) && memcmp(pos, marker, strlen(marker)) == 0) { found = pos; /* keep the last occurrence */ } pos++; } if (!found) { free(buf); return NULL; } /* After "sep_meta" the protobuf stores the value string. Skip past the key and any protobuf length bytes until '{' */ char* json_start = found + strlen(marker); while (json_start < buf + n && *json_start != '{') json_start++; if (json_start >= buf + n) { free(buf); return NULL; } /* Find matching closing '}' */ int depth = 0; char* p = json_start; while (p < buf + n) { if (*p == '{') depth++; else if (*p == '}') { depth--; if (depth == 0) break; } p++; } if (depth != 0) { free(buf); return NULL; } size_t json_len = (size_t)(p - json_start) + 1; char* result = (char*)malloc(json_len + 1); memcpy(result, json_start, json_len); result[json_len] = '\0'; free(buf); return result; } /* Example usage */ int main(void) { char* meta = read_onnx_sep_meta("Reverb_HQ_By_FoxJoy.onnx"); if (meta) { printf("%s\n", meta); free(meta); } return 0; } ``` --- ## 🔬 `sep_meta` JSON Schema The `sep_meta` value is a compact JSON object. Fields differ by architecture: ### MDX-Net models ```json { "arch": "MDX", "primary_stem": "Vocals", "secondary_stem": "Instrumental", "sample_rate": 44100, "n_fft": 7680, "hop_length": 1024, "dim_f": 3072, "dim_t": 256, "compensate": 1.021, "overlap": 0.25 } ``` | Field | Type | Description | |-------|------|-------------| | `n_fft` | int | STFT window size | | `hop_length` | int | STFT hop length | | `dim_f` | int | Frequency bins fed to the network | | `dim_t` | int | Time frames per chunk | | `compensate` | float | Amplitude gain applied after iSTFT | | `overlap` | float | Overlap ratio between consecutive chunks (0–0.75) | **ONNX I/O:** - Input: `input` — shape `(1, 4, dim_f, dim_t)` — four channels: `[real_L, imag_L, real_R, imag_R]` - Output: `output` — same shape as input ### VR Architecture models ```json { "arch": "VR", "primary_stem": "No Reverb", "secondary_stem": "Reverb", "sample_rate": 44100, "vr_model_param": "4band_v3", "bins": 672, "window_size": 512, "is_vr51": true, "nn_arch_size": 218, "model_capacity": [32, 128], "band_params": { "1": {"sr": 11025, "n_fft": 2048, "trim": 0, "window": "hann", "agg": 10}, "2": {"sr": 22050, "n_fft": 2048, "trim": 48, "window": "hann", "agg": 10}, "3": {"sr": 44100, "n_fft": 2048, "trim": 48, "window": "hann", "agg": 10}, "4": {"sr": 44100, "n_fft": 8192, "trim": 0, "window": "hann", "agg": 10} } } ``` | Field | Type | Description | |-------|------|-------------| | `vr_model_param` | string | Band config preset name | | `bins` | int | Total frequency bins after multi-band STFT | | `window_size` | int | Window size for the primary STFT pass | | `is_vr51` | bool | `true` = CascadedNet (UVR v5.1), `false` = v4 | | `nn_arch_size` | int | Network depth parameter | | `band_params` | object | Per-band STFT config used to build the input spectrogram | **ONNX I/O:** - Input: `input` — shape `(batch, 2, bins, frames)` — stereo multi-band magnitude spectrogram - Output: `output` — shape `(batch, 2, bins, frames)` — predicted source mask ### Roformer models (MelBandRoformer / BSRoformer) > **Note:** Roformer ONNX exports contain only the **transformer core**. STFT and iSTFT are not part of the ONNX graph (complex ops are not exportable). The inference wrapper must perform STFT → band-split → ONNX → scatter → iSTFT around the model. ```json { "arch": "ROFORMER", "roformer_type": "MelBandRoformer", "primary_stem": "No Reverb", "secondary_stem": "Reverb", "sample_rate": 44100, "chunk_size": 264600, "hop_length": 1024, "n_fft": 2048, "dim_t": 256, "num_bands": 64, "dim": 128, "frames": 256, "n_freq_indices": 384, "num_bins": 1025, "n_full_freqs": 1025, "overlap": 2, "freq_indices": [...], "num_bands_per_freq": [...] } ``` | Field | Type | Description | |-------|------|-------------| | `chunk_size` | int | Audio samples per inference chunk | | `hop_length` / `n_fft` | int | STFT params applied **outside** the ONNX graph | | `num_bands` | int | Number of mel bands | | `frames` | int | Time frames (= `dim_t`) | | `n_freq_indices` | int | Number of selected frequency bins after mel gather | | `freq_indices` | int[] | Frequency bin indices to gather from full STFT before passing to network | | `num_bands_per_freq` | int[] | Band assignment per frequency bin (for scatter-add in iSTFT) | | `overlap` | int | Number of overlapping inference passes per chunk | **ONNX I/O:** - Input: `x_bands` — shape `(1, frames, num_bands, dim)` — band-split STFT features - Output: `masks` — shape `(1, 1, n_freq_indices, frames, 2)` — complex mask --- ## 🚀 Quickstart with `audio-separator` The recommended way to run these models is [`audio-separator`](https://github.com/nomadkaraoke/python-audio-separator) — the same inference engine that powers [Ultimate Vocal Remover](https://github.com/Anjok07/ultimatevocalremovergui). All `.onnx` and `.pth` files in this repo are fully compatible with `audio-separator` out of the box. ### Installation ```bash # CPU only (macOS / Linux / Windows) pip install "audio-separator[cpu]" # Nvidia GPU (CUDA) pip install "audio-separator[gpu]" # Apple Silicon (CoreML acceleration) pip install "audio-separator[cpu]" # CoreML is auto-detected on macOS ``` ### CLI usage Download the model file from the **Files and versions** tab, then point `--model_file_dir` at the folder containing it: ```bash # Remove reverb from a mix audio-separator mix.wav \ --model_filename Reverb_HQ_By_FoxJoy.onnx \ --model_file_dir /path/to/downloaded/models \ --output_dir ./output \ --output_format WAV # Vocal / instrumental separation audio-separator mix.wav \ --model_filename UVR-MDX-NET-Inst_HQ_3.onnx \ --model_file_dir /path/to/downloaded/models \ --output_dir ./output ``` ### Python API ```python from audio_separator.separator import Separator separator = Separator( model_file_dir="/path/to/downloaded/models", # folder with the .onnx files output_dir="./output", output_format="WAV", ) # Load and run — arch is auto-detected from the file extension separator.load_model("Reverb_HQ_By_FoxJoy.onnx") output_files = separator.separate("mix.wav") print("Output:", output_files) ``` ### Adjusting inference parameters `audio-separator` exposes the most important per-arch parameters: ```python # MDX — tune overlap and segment size separator = Separator( model_file_dir="/path/to/downloaded/models", mdx_params={ "hop_length": 1024, "segment_size": 256, # dim_t from sep_meta "overlap": 0.25, # matches sep_meta default "batch_size": 1, "enable_denoise": False, } ) # VR — tune aggression and window size separator = Separator( model_file_dir="/path/to/downloaded/models", vr_params={ "batch_size": 1, "window_size": 512, # 320 = slower but better "aggression": 5, # 0-100, higher = more aggressive separation "enable_tta": False, "high_end_process": False, } ) separator.load_model("UVR-DeEcho-DeReverb.onnx") output_files = separator.separate("mix.wav") ``` > **Note on `sep_meta`:** `audio-separator` reads model parameters from its own internal `mdx_model_data.json` lookup table — it does **not** read the embedded `sep_meta` metadata. The `sep_meta` field is provided for custom inference pipelines and other runtimes (see sections below). embed_sep_meta(onnx_path: str, meta: dict) -> None: ```python model = onnx.load(onnx_path) # Remove stale entry if present existing = [p for p in model.metadata_props if p.key != "sep_meta"] del model.metadata_props[:] model.metadata_props.extend(existing) # Add new entry entry = model.metadata_props.add() entry.key = "sep_meta" entry.value = json.dumps(meta, separators=(",", ":")) onnx.save_model(model, onnx_path) ``` # Example — MDX model ```python embed_sep_meta("my_model.onnx", { "arch": "MDX", "primary_stem": "Vocals", "secondary_stem": "Instrumental", "sample_rate": 44100, "n_fft": 7680, "hop_length": 1024, "dim_f": 3072, "dim_t": 256, "compensate": 1.021, "overlap": 0.25, }) ``` --- ## 📋 Requirements ``` onnxruntime >= 1.16 numpy >= 1.24 soundfile >= 0.12 ``` For metadata-only access (no inference): ``` onnx >= 1.14 ``` ## 📄 License The repository structure and conversion scripts are licensed under the [MIT License](https://www.google.com/search?q=LICENSE). *Please check the original model licenses before using them for commercial purposes.* --- ## 🙏 Acknowledgments & Credits Special thanks to the authors and maintainers of: * [Ultimate Vocal Remover (UVR5)](https://www.google.com/search?q=https://github.com/Anemone95/Ultimate-Vocal-Remover-GUI) * [audio-separator](https://www.google.com/search?q=https://github.com/beverlis/audio-separator) * [ONNX Runtime](https://onnxruntime.ai/)