AudioSeparatorONNX / README.md
raydotac's picture
Update README.md
c02cba6 verified
|
Raw History Blame
12.7 kB
---
license: mit
pipeline_tag: audio-to-audio
tags:
- audio-separation
- vocal-remover
- stem-separation
- onnx
- onnxruntime
---
# AudioSeparatorONNX
Collection of popular audio separation and vocal removal models converted to **ONNX** format, optimized for high-performance CPU/GPU inference using `onnxruntime`.
## πŸ“Œ Overview
This repository provides pre-converted ONNX versions of popular music source separation (MSS) and vocal removal models (such as UVR5 / Ultimate Vocal Remover models, MDX-Net, VR Architecture, Demucs, etc.).
Using ONNX models allows for:
- **Cross-platform compatibility** (Python, C++, C#, Rust, JS, etc.).
- **Faster execution** with `onnxruntime` using CPU, CUDA, DirectML, CoreML, or TensorRT.
- **Lightweight deployments** without requiring heavy frameworks like PyTorch or TensorFlow in production.
---
## πŸ’» Quickstart
## πŸ“¦ Embedded Metadata (`sep_meta`)
Every `.onnx` file in this repository is **self-contained** β€” no separate sidecar `.json` file is needed. All inference parameters are embedded directly inside the ONNX model as a `metadata_props` entry with key `sep_meta`.
### Reading metadata with Python (`onnxruntime`)
```python
import json
import onnxruntime as ort
sess = ort.InferenceSession("Reverb_HQ_By_FoxJoy.onnx", providers=["CPUExecutionProvider"])
meta_map = sess.get_modelmeta().custom_metadata_map # dict[str, str]
sep_meta = json.loads(meta_map["sep_meta"])
print(sep_meta["arch"]) # "MDX" | "VR" | "ROFORMER"
print(sep_meta["primary_stem"]) # "No Reverb"
print(sep_meta["secondary_stem"]) # "Reverb"
```
### Reading metadata with Python (`onnx` library, no ORT session)
```python
import json
import onnx
model = onnx.load("UVR-MDX-NET-Inst_HQ_1.onnx")
meta_map = {p.key: p.value for p in model.metadata_props}
sep_meta = json.loads(meta_map["sep_meta"])
```
### Reading metadata with C (fast byte-scan, no ORT dependency)
The `sep_meta` key is stored near the **end** of the ONNX protobuf binary, after all weight initializers. You can extract it with a simple byte scan β€” no need to load the entire model into memory:
```c
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
/* Returns a heap-allocated JSON string, or NULL on failure.
Caller must free() the result. */
char* read_onnx_sep_meta(const char* onnx_path) {
FILE* f = fopen(onnx_path, "rb");
if (!f) return NULL;
/* Read the last 64 KB β€” sep_meta is always at the end */
const size_t SCAN_SIZE = 64 * 1024;
fseek(f, 0, SEEK_END);
long file_size = ftell(f);
size_t read_size = (file_size < (long)SCAN_SIZE) ? (size_t)file_size : SCAN_SIZE;
long offset = file_size - (long)read_size;
fseek(f, offset, SEEK_SET);
char* buf = (char*)malloc(read_size + 1);
if (!buf) { fclose(f); return NULL; }
size_t n = fread(buf, 1, read_size, f);
fclose(f);
buf[n] = '\0';
/* Find the marker string "sep_meta" */
const char* marker = "sep_meta";
char* pos = buf;
char* found = NULL;
while ((pos = (char*)memchr(pos, 's', buf + n - pos)) != NULL) {
if ((size_t)(buf + n - pos) >= strlen(marker) &&
memcmp(pos, marker, strlen(marker)) == 0) {
found = pos; /* keep the last occurrence */
}
pos++;
}
if (!found) { free(buf); return NULL; }
/* After "sep_meta" the protobuf stores the value string.
Skip past the key and any protobuf length bytes until '{' */
char* json_start = found + strlen(marker);
while (json_start < buf + n && *json_start != '{') json_start++;
if (json_start >= buf + n) { free(buf); return NULL; }
/* Find matching closing '}' */
int depth = 0;
char* p = json_start;
while (p < buf + n) {
if (*p == '{') depth++;
else if (*p == '}') { depth--; if (depth == 0) break; }
p++;
}
if (depth != 0) { free(buf); return NULL; }
size_t json_len = (size_t)(p - json_start) + 1;
char* result = (char*)malloc(json_len + 1);
memcpy(result, json_start, json_len);
result[json_len] = '\0';
free(buf);
return result;
}
/* Example usage */
int main(void) {
char* meta = read_onnx_sep_meta("Reverb_HQ_By_FoxJoy.onnx");
if (meta) {
printf("%s\n", meta);
free(meta);
}
return 0;
}
```
---
## πŸ”¬ `sep_meta` JSON Schema
The `sep_meta` value is a compact JSON object. Fields differ by architecture:
### MDX-Net models
```json
{
"arch": "MDX",
"primary_stem": "Vocals",
"secondary_stem": "Instrumental",
"sample_rate": 44100,
"n_fft": 7680,
"hop_length": 1024,
"dim_f": 3072,
"dim_t": 256,
"compensate": 1.021,
"overlap": 0.25
}
```
| Field | Type | Description |
|-------|------|-------------|
| `n_fft` | int | STFT window size |
| `hop_length` | int | STFT hop length |
| `dim_f` | int | Frequency bins fed to the network |
| `dim_t` | int | Time frames per chunk |
| `compensate` | float | Amplitude gain applied after iSTFT |
| `overlap` | float | Overlap ratio between consecutive chunks (0–0.75) |
**ONNX I/O:**
- Input: `input` β€” shape `(1, 4, dim_f, dim_t)` β€” four channels: `[real_L, imag_L, real_R, imag_R]`
- Output: `output` β€” same shape as input
### VR Architecture models
```json
{
"arch": "VR",
"primary_stem": "No Reverb",
"secondary_stem": "Reverb",
"sample_rate": 44100,
"vr_model_param": "4band_v3",
"bins": 672,
"window_size": 512,
"is_vr51": true,
"nn_arch_size": 218,
"model_capacity": [32, 128],
"band_params": {
"1": {"sr": 11025, "n_fft": 2048, "trim": 0, "window": "hann", "agg": 10},
"2": {"sr": 22050, "n_fft": 2048, "trim": 48, "window": "hann", "agg": 10},
"3": {"sr": 44100, "n_fft": 2048, "trim": 48, "window": "hann", "agg": 10},
"4": {"sr": 44100, "n_fft": 8192, "trim": 0, "window": "hann", "agg": 10}
}
}
```
| Field | Type | Description |
|-------|------|-------------|
| `vr_model_param` | string | Band config preset name |
| `bins` | int | Total frequency bins after multi-band STFT |
| `window_size` | int | Window size for the primary STFT pass |
| `is_vr51` | bool | `true` = CascadedNet (UVR v5.1), `false` = v4 |
| `nn_arch_size` | int | Network depth parameter |
| `band_params` | object | Per-band STFT config used to build the input spectrogram |
**ONNX I/O:**
- Input: `input` β€” shape `(batch, 2, bins, frames)` β€” stereo multi-band magnitude spectrogram
- Output: `output` β€” shape `(batch, 2, bins, frames)` β€” predicted source mask
### Roformer models (MelBandRoformer / BSRoformer)
> **Note:** Roformer ONNX exports contain only the **transformer core**. STFT and iSTFT are not part of the ONNX graph (complex ops are not exportable). The inference wrapper must perform STFT β†’ band-split β†’ ONNX β†’ scatter β†’ iSTFT around the model.
```json
{
"arch": "ROFORMER",
"roformer_type": "MelBandRoformer",
"primary_stem": "No Reverb",
"secondary_stem": "Reverb",
"sample_rate": 44100,
"chunk_size": 264600,
"hop_length": 1024,
"n_fft": 2048,
"dim_t": 256,
"num_bands": 64,
"dim": 128,
"frames": 256,
"n_freq_indices": 384,
"num_bins": 1025,
"n_full_freqs": 1025,
"overlap": 2,
"freq_indices": [...],
"num_bands_per_freq": [...]
}
```
| Field | Type | Description |
|-------|------|-------------|
| `chunk_size` | int | Audio samples per inference chunk |
| `hop_length` / `n_fft` | int | STFT params applied **outside** the ONNX graph |
| `num_bands` | int | Number of mel bands |
| `frames` | int | Time frames (= `dim_t`) |
| `n_freq_indices` | int | Number of selected frequency bins after mel gather |
| `freq_indices` | int[] | Frequency bin indices to gather from full STFT before passing to network |
| `num_bands_per_freq` | int[] | Band assignment per frequency bin (for scatter-add in iSTFT) |
| `overlap` | int | Number of overlapping inference passes per chunk |
**ONNX I/O:**
- Input: `x_bands` β€” shape `(1, frames, num_bands, dim)` β€” band-split STFT features
- Output: `masks` β€” shape `(1, 1, n_freq_indices, frames, 2)` β€” complex mask
---
## πŸš€ Quickstart with `audio-separator`
The recommended way to run these models is [`audio-separator`](https://github.com/nomadkaraoke/python-audio-separator) β€” the same inference engine that powers [Ultimate Vocal Remover](https://github.com/Anjok07/ultimatevocalremovergui).
All `.onnx` and `.pth` files in this repo are fully compatible with `audio-separator` out of the box.
### Installation
```bash
# CPU only (macOS / Linux / Windows)
pip install "audio-separator[cpu]"
# Nvidia GPU (CUDA)
pip install "audio-separator[gpu]"
# Apple Silicon (CoreML acceleration)
pip install "audio-separator[cpu]" # CoreML is auto-detected on macOS
```
### CLI usage
Download the model file from the **Files and versions** tab, then point `--model_file_dir` at the folder containing it:
```bash
# Remove reverb from a mix
audio-separator mix.wav \
--model_filename Reverb_HQ_By_FoxJoy.onnx \
--model_file_dir /path/to/downloaded/models \
--output_dir ./output \
--output_format WAV
# Vocal / instrumental separation
audio-separator mix.wav \
--model_filename UVR-MDX-NET-Inst_HQ_3.onnx \
--model_file_dir /path/to/downloaded/models \
--output_dir ./output
```
### Python API
```python
from audio_separator.separator import Separator
separator = Separator(
model_file_dir="/path/to/downloaded/models", # folder with the .onnx files
output_dir="./output",
output_format="WAV",
)
# Load and run β€” arch is auto-detected from the file extension
separator.load_model("Reverb_HQ_By_FoxJoy.onnx")
output_files = separator.separate("mix.wav")
print("Output:", output_files)
```
### Adjusting inference parameters
`audio-separator` exposes the most important per-arch parameters:
```python
# MDX β€” tune overlap and segment size
separator = Separator(
model_file_dir="/path/to/downloaded/models",
mdx_params={
"hop_length": 1024,
"segment_size": 256, # dim_t from sep_meta
"overlap": 0.25, # matches sep_meta default
"batch_size": 1,
"enable_denoise": False,
}
)
# VR β€” tune aggression and window size
separator = Separator(
model_file_dir="/path/to/downloaded/models",
vr_params={
"batch_size": 1,
"window_size": 512, # 320 = slower but better
"aggression": 5, # 0-100, higher = more aggressive separation
"enable_tta": False,
"high_end_process": False,
}
)
separator.load_model("UVR-DeEcho-DeReverb.onnx")
output_files = separator.separate("mix.wav")
```
> **Note on `sep_meta`:** `audio-separator` reads model parameters from its own internal `mdx_model_data.json` lookup table β€” it does **not** read the embedded `sep_meta` metadata. The `sep_meta` field is provided for custom inference pipelines and other runtimes (see sections below). embed_sep_meta(onnx_path: str, meta: dict) -> None:
```python
model = onnx.load(onnx_path)
# Remove stale entry if present
existing = [p for p in model.metadata_props if p.key != "sep_meta"]
del model.metadata_props[:]
model.metadata_props.extend(existing)
# Add new entry
entry = model.metadata_props.add()
entry.key = "sep_meta"
entry.value = json.dumps(meta, separators=(",", ":"))
onnx.save_model(model, onnx_path)
```
# Example β€” MDX model
```python
embed_sep_meta("my_model.onnx", {
"arch": "MDX",
"primary_stem": "Vocals",
"secondary_stem": "Instrumental",
"sample_rate": 44100,
"n_fft": 7680,
"hop_length": 1024,
"dim_f": 3072,
"dim_t": 256,
"compensate": 1.021,
"overlap": 0.25,
})
```
---
## πŸ“‹ Requirements
```
onnxruntime >= 1.16
numpy >= 1.24
soundfile >= 0.12
```
For metadata-only access (no inference):
```
onnx >= 1.14
```
## πŸ“„ License
The repository structure and conversion scripts are licensed under the [MIT License](https://www.google.com/search?q=LICENSE).
*Please check the original model licenses before using them for commercial purposes.*
---
## πŸ™ Acknowledgments & Credits
Special thanks to the authors and maintainers of:
* [Ultimate Vocal Remover (UVR5)](https://www.google.com/search?q=https://github.com/Anemone95/Ultimate-Vocal-Remover-GUI)
* [audio-separator](https://www.google.com/search?q=https://github.com/beverlis/audio-separator)
* [ONNX Runtime](https://onnxruntime.ai/)