|
Download README.md from raydotac/AudioSeparatorONNX: direct link, hf CLI and curl.
- Browser
- Download file 12.7 kB
-
https://huggingface.co/raydotac/AudioSeparatorONNX/resolve/ce83069c9ba90dd8d91a86d55c7107f1e8ee2d8c/README.md
- Command line
-
hf download hf://raydotac/AudioSeparatorONNX@ce83069c9ba90dd8d91a86d55c7107f1e8ee2d8c/README.md
-
curl -L -o README.md https://huggingface.co/raydotac/AudioSeparatorONNX/resolve/ce83069c9ba90dd8d91a86d55c7107f1e8ee2d8c/README.md
12.7 kB
| license: mit | |
| pipeline_tag: audio-to-audio | |
| tags: | |
| - audio-separation | |
| - vocal-remover | |
| - stem-separation | |
| - onnx | |
| - onnxruntime | |
| # AudioSeparatorONNX | |
| Collection of popular audio separation and vocal removal models converted to **ONNX** format, optimized for high-performance CPU/GPU inference using `onnxruntime`. | |
| ## π Overview | |
| This repository provides pre-converted ONNX versions of popular music source separation (MSS) and vocal removal models (such as UVR5 / Ultimate Vocal Remover models, MDX-Net, VR Architecture, Demucs, etc.). | |
| Using ONNX models allows for: | |
| - **Cross-platform compatibility** (Python, C++, C#, Rust, JS, etc.). | |
| - **Faster execution** with `onnxruntime` using CPU, CUDA, DirectML, CoreML, or TensorRT. | |
| - **Lightweight deployments** without requiring heavy frameworks like PyTorch or TensorFlow in production. | |
| --- | |
| ## π» Quickstart | |
| ## π¦ Embedded Metadata (`sep_meta`) | |
| Every `.onnx` file in this repository is **self-contained** β no separate sidecar `.json` file is needed. All inference parameters are embedded directly inside the ONNX model as a `metadata_props` entry with key `sep_meta`. | |
| ### Reading metadata with Python (`onnxruntime`) | |
| ```python | |
| import json | |
| import onnxruntime as ort | |
| sess = ort.InferenceSession("Reverb_HQ_By_FoxJoy.onnx", providers=["CPUExecutionProvider"]) | |
| meta_map = sess.get_modelmeta().custom_metadata_map # dict[str, str] | |
| sep_meta = json.loads(meta_map["sep_meta"]) | |
| print(sep_meta["arch"]) # "MDX" | "VR" | "ROFORMER" | |
| print(sep_meta["primary_stem"]) # "No Reverb" | |
| print(sep_meta["secondary_stem"]) # "Reverb" | |
| ``` | |
| ### Reading metadata with Python (`onnx` library, no ORT session) | |
| ```python | |
| import json | |
| import onnx | |
| model = onnx.load("UVR-MDX-NET-Inst_HQ_1.onnx") | |
| meta_map = {p.key: p.value for p in model.metadata_props} | |
| sep_meta = json.loads(meta_map["sep_meta"]) | |
| ``` | |
| ### Reading metadata with C (fast byte-scan, no ORT dependency) | |
| The `sep_meta` key is stored near the **end** of the ONNX protobuf binary, after all weight initializers. You can extract it with a simple byte scan β no need to load the entire model into memory: | |
| ```c | |
| #include <stdio.h> | |
| #include <stdlib.h> | |
| #include <string.h> | |
| /* Returns a heap-allocated JSON string, or NULL on failure. | |
| Caller must free() the result. */ | |
| char* read_onnx_sep_meta(const char* onnx_path) { | |
| FILE* f = fopen(onnx_path, "rb"); | |
| if (!f) return NULL; | |
| /* Read the last 64 KB β sep_meta is always at the end */ | |
| const size_t SCAN_SIZE = 64 * 1024; | |
| fseek(f, 0, SEEK_END); | |
| long file_size = ftell(f); | |
| size_t read_size = (file_size < (long)SCAN_SIZE) ? (size_t)file_size : SCAN_SIZE; | |
| long offset = file_size - (long)read_size; | |
| fseek(f, offset, SEEK_SET); | |
| char* buf = (char*)malloc(read_size + 1); | |
| if (!buf) { fclose(f); return NULL; } | |
| size_t n = fread(buf, 1, read_size, f); | |
| fclose(f); | |
| buf[n] = '\0'; | |
| /* Find the marker string "sep_meta" */ | |
| const char* marker = "sep_meta"; | |
| char* pos = buf; | |
| char* found = NULL; | |
| while ((pos = (char*)memchr(pos, 's', buf + n - pos)) != NULL) { | |
| if ((size_t)(buf + n - pos) >= strlen(marker) && | |
| memcmp(pos, marker, strlen(marker)) == 0) { | |
| found = pos; /* keep the last occurrence */ | |
| } | |
| pos++; | |
| } | |
| if (!found) { free(buf); return NULL; } | |
| /* After "sep_meta" the protobuf stores the value string. | |
| Skip past the key and any protobuf length bytes until '{' */ | |
| char* json_start = found + strlen(marker); | |
| while (json_start < buf + n && *json_start != '{') json_start++; | |
| if (json_start >= buf + n) { free(buf); return NULL; } | |
| /* Find matching closing '}' */ | |
| int depth = 0; | |
| char* p = json_start; | |
| while (p < buf + n) { | |
| if (*p == '{') depth++; | |
| else if (*p == '}') { depth--; if (depth == 0) break; } | |
| p++; | |
| } | |
| if (depth != 0) { free(buf); return NULL; } | |
| size_t json_len = (size_t)(p - json_start) + 1; | |
| char* result = (char*)malloc(json_len + 1); | |
| memcpy(result, json_start, json_len); | |
| result[json_len] = '\0'; | |
| free(buf); | |
| return result; | |
| } | |
| /* Example usage */ | |
| int main(void) { | |
| char* meta = read_onnx_sep_meta("Reverb_HQ_By_FoxJoy.onnx"); | |
| if (meta) { | |
| printf("%s\n", meta); | |
| free(meta); | |
| } | |
| return 0; | |
| } | |
| ``` | |
| --- | |
| ## π¬ `sep_meta` JSON Schema | |
| The `sep_meta` value is a compact JSON object. Fields differ by architecture: | |
| ### MDX-Net models | |
| ```json | |
| { | |
| "arch": "MDX", | |
| "primary_stem": "Vocals", | |
| "secondary_stem": "Instrumental", | |
| "sample_rate": 44100, | |
| "n_fft": 7680, | |
| "hop_length": 1024, | |
| "dim_f": 3072, | |
| "dim_t": 256, | |
| "compensate": 1.021, | |
| "overlap": 0.25 | |
| } | |
| ``` | |
| | Field | Type | Description | | |
| |-------|------|-------------| | |
| | `n_fft` | int | STFT window size | | |
| | `hop_length` | int | STFT hop length | | |
| | `dim_f` | int | Frequency bins fed to the network | | |
| | `dim_t` | int | Time frames per chunk | | |
| | `compensate` | float | Amplitude gain applied after iSTFT | | |
| | `overlap` | float | Overlap ratio between consecutive chunks (0β0.75) | | |
| **ONNX I/O:** | |
| - Input: `input` β shape `(1, 4, dim_f, dim_t)` β four channels: `[real_L, imag_L, real_R, imag_R]` | |
| - Output: `output` β same shape as input | |
| ### VR Architecture models | |
| ```json | |
| { | |
| "arch": "VR", | |
| "primary_stem": "No Reverb", | |
| "secondary_stem": "Reverb", | |
| "sample_rate": 44100, | |
| "vr_model_param": "4band_v3", | |
| "bins": 672, | |
| "window_size": 512, | |
| "is_vr51": true, | |
| "nn_arch_size": 218, | |
| "model_capacity": [32, 128], | |
| "band_params": { | |
| "1": {"sr": 11025, "n_fft": 2048, "trim": 0, "window": "hann", "agg": 10}, | |
| "2": {"sr": 22050, "n_fft": 2048, "trim": 48, "window": "hann", "agg": 10}, | |
| "3": {"sr": 44100, "n_fft": 2048, "trim": 48, "window": "hann", "agg": 10}, | |
| "4": {"sr": 44100, "n_fft": 8192, "trim": 0, "window": "hann", "agg": 10} | |
| } | |
| } | |
| ``` | |
| | Field | Type | Description | | |
| |-------|------|-------------| | |
| | `vr_model_param` | string | Band config preset name | | |
| | `bins` | int | Total frequency bins after multi-band STFT | | |
| | `window_size` | int | Window size for the primary STFT pass | | |
| | `is_vr51` | bool | `true` = CascadedNet (UVR v5.1), `false` = v4 | | |
| | `nn_arch_size` | int | Network depth parameter | | |
| | `band_params` | object | Per-band STFT config used to build the input spectrogram | | |
| **ONNX I/O:** | |
| - Input: `input` β shape `(batch, 2, bins, frames)` β stereo multi-band magnitude spectrogram | |
| - Output: `output` β shape `(batch, 2, bins, frames)` β predicted source mask | |
| ### Roformer models (MelBandRoformer / BSRoformer) | |
| > **Note:** Roformer ONNX exports contain only the **transformer core**. STFT and iSTFT are not part of the ONNX graph (complex ops are not exportable). The inference wrapper must perform STFT β band-split β ONNX β scatter β iSTFT around the model. | |
| ```json | |
| { | |
| "arch": "ROFORMER", | |
| "roformer_type": "MelBandRoformer", | |
| "primary_stem": "No Reverb", | |
| "secondary_stem": "Reverb", | |
| "sample_rate": 44100, | |
| "chunk_size": 264600, | |
| "hop_length": 1024, | |
| "n_fft": 2048, | |
| "dim_t": 256, | |
| "num_bands": 64, | |
| "dim": 128, | |
| "frames": 256, | |
| "n_freq_indices": 384, | |
| "num_bins": 1025, | |
| "n_full_freqs": 1025, | |
| "overlap": 2, | |
| "freq_indices": [...], | |
| "num_bands_per_freq": [...] | |
| } | |
| ``` | |
| | Field | Type | Description | | |
| |-------|------|-------------| | |
| | `chunk_size` | int | Audio samples per inference chunk | | |
| | `hop_length` / `n_fft` | int | STFT params applied **outside** the ONNX graph | | |
| | `num_bands` | int | Number of mel bands | | |
| | `frames` | int | Time frames (= `dim_t`) | | |
| | `n_freq_indices` | int | Number of selected frequency bins after mel gather | | |
| | `freq_indices` | int[] | Frequency bin indices to gather from full STFT before passing to network | | |
| | `num_bands_per_freq` | int[] | Band assignment per frequency bin (for scatter-add in iSTFT) | | |
| | `overlap` | int | Number of overlapping inference passes per chunk | | |
| **ONNX I/O:** | |
| - Input: `x_bands` β shape `(1, frames, num_bands, dim)` β band-split STFT features | |
| - Output: `masks` β shape `(1, 1, n_freq_indices, frames, 2)` β complex mask | |
| --- | |
| ## π Quickstart with `audio-separator` | |
| The recommended way to run these models is [`audio-separator`](https://github.com/nomadkaraoke/python-audio-separator) β the same inference engine that powers [Ultimate Vocal Remover](https://github.com/Anjok07/ultimatevocalremovergui). | |
| All `.onnx` and `.pth` files in this repo are fully compatible with `audio-separator` out of the box. | |
| ### Installation | |
| ```bash | |
| # CPU only (macOS / Linux / Windows) | |
| pip install "audio-separator[cpu]" | |
| # Nvidia GPU (CUDA) | |
| pip install "audio-separator[gpu]" | |
| # Apple Silicon (CoreML acceleration) | |
| pip install "audio-separator[cpu]" # CoreML is auto-detected on macOS | |
| ``` | |
| ### CLI usage | |
| Download the model file from the **Files and versions** tab, then point `--model_file_dir` at the folder containing it: | |
| ```bash | |
| # Remove reverb from a mix | |
| audio-separator mix.wav \ | |
| --model_filename Reverb_HQ_By_FoxJoy.onnx \ | |
| --model_file_dir /path/to/downloaded/models \ | |
| --output_dir ./output \ | |
| --output_format WAV | |
| # Vocal / instrumental separation | |
| audio-separator mix.wav \ | |
| --model_filename UVR-MDX-NET-Inst_HQ_3.onnx \ | |
| --model_file_dir /path/to/downloaded/models \ | |
| --output_dir ./output | |
| ``` | |
| ### Python API | |
| ```python | |
| from audio_separator.separator import Separator | |
| separator = Separator( | |
| model_file_dir="/path/to/downloaded/models", # folder with the .onnx files | |
| output_dir="./output", | |
| output_format="WAV", | |
| ) | |
| # Load and run β arch is auto-detected from the file extension | |
| separator.load_model("Reverb_HQ_By_FoxJoy.onnx") | |
| output_files = separator.separate("mix.wav") | |
| print("Output:", output_files) | |
| ``` | |
| ### Adjusting inference parameters | |
| `audio-separator` exposes the most important per-arch parameters: | |
| ```python | |
| # MDX β tune overlap and segment size | |
| separator = Separator( | |
| model_file_dir="/path/to/downloaded/models", | |
| mdx_params={ | |
| "hop_length": 1024, | |
| "segment_size": 256, # dim_t from sep_meta | |
| "overlap": 0.25, # matches sep_meta default | |
| "batch_size": 1, | |
| "enable_denoise": False, | |
| } | |
| ) | |
| # VR β tune aggression and window size | |
| separator = Separator( | |
| model_file_dir="/path/to/downloaded/models", | |
| vr_params={ | |
| "batch_size": 1, | |
| "window_size": 512, # 320 = slower but better | |
| "aggression": 5, # 0-100, higher = more aggressive separation | |
| "enable_tta": False, | |
| "high_end_process": False, | |
| } | |
| ) | |
| separator.load_model("UVR-DeEcho-DeReverb.onnx") | |
| output_files = separator.separate("mix.wav") | |
| ``` | |
| > **Note on `sep_meta`:** `audio-separator` reads model parameters from its own internal `mdx_model_data.json` lookup table β it does **not** read the embedded `sep_meta` metadata. The `sep_meta` field is provided for custom inference pipelines and other runtimes (see sections below). embed_sep_meta(onnx_path: str, meta: dict) -> None: | |
| ```python | |
| model = onnx.load(onnx_path) | |
| # Remove stale entry if present | |
| existing = [p for p in model.metadata_props if p.key != "sep_meta"] | |
| del model.metadata_props[:] | |
| model.metadata_props.extend(existing) | |
| # Add new entry | |
| entry = model.metadata_props.add() | |
| entry.key = "sep_meta" | |
| entry.value = json.dumps(meta, separators=(",", ":")) | |
| onnx.save_model(model, onnx_path) | |
| ``` | |
| # Example β MDX model | |
| ```python | |
| embed_sep_meta("my_model.onnx", { | |
| "arch": "MDX", | |
| "primary_stem": "Vocals", | |
| "secondary_stem": "Instrumental", | |
| "sample_rate": 44100, | |
| "n_fft": 7680, | |
| "hop_length": 1024, | |
| "dim_f": 3072, | |
| "dim_t": 256, | |
| "compensate": 1.021, | |
| "overlap": 0.25, | |
| }) | |
| ``` | |
| --- | |
| ## π Requirements | |
| ``` | |
| onnxruntime >= 1.16 | |
| numpy >= 1.24 | |
| soundfile >= 0.12 | |
| ``` | |
| For metadata-only access (no inference): | |
| ``` | |
| onnx >= 1.14 | |
| ``` | |
| ## π License | |
| The repository structure and conversion scripts are licensed under the [MIT License](https://www.google.com/search?q=LICENSE). | |
| *Please check the original model licenses before using them for commercial purposes.* | |
| --- | |
| ## π Acknowledgments & Credits | |
| Special thanks to the authors and maintainers of: | |
| * [Ultimate Vocal Remover (UVR5)](https://www.google.com/search?q=https://github.com/Anemone95/Ultimate-Vocal-Remover-GUI) | |
| * [audio-separator](https://www.google.com/search?q=https://github.com/beverlis/audio-separator) | |
| * [ONNX Runtime](https://onnxruntime.ai/) | |