Download README.md from raydotac/AudioSeparatorONNX: direct link, hf CLI and curl.
- Browser
- Download file 12.7 kB
-
https://huggingface.co/raydotac/AudioSeparatorONNX/resolve/ce83069c9ba90dd8d91a86d55c7107f1e8ee2d8c/README.md
- Command line
-
hf download hf://raydotac/AudioSeparatorONNX@ce83069c9ba90dd8d91a86d55c7107f1e8ee2d8c/README.md
-
curl -L -o README.md https://huggingface.co/raydotac/AudioSeparatorONNX/resolve/ce83069c9ba90dd8d91a86d55c7107f1e8ee2d8c/README.md
license: mit
pipeline_tag: audio-to-audio
tags:
- audio-separation
- vocal-remover
- stem-separation
- onnx
- onnxruntime
AudioSeparatorONNX
Collection of popular audio separation and vocal removal models converted to ONNX format, optimized for high-performance CPU/GPU inference using onnxruntime.
π Overview
This repository provides pre-converted ONNX versions of popular music source separation (MSS) and vocal removal models (such as UVR5 / Ultimate Vocal Remover models, MDX-Net, VR Architecture, Demucs, etc.).
Using ONNX models allows for:
- Cross-platform compatibility (Python, C++, C#, Rust, JS, etc.).
- Faster execution with
onnxruntimeusing CPU, CUDA, DirectML, CoreML, or TensorRT. - Lightweight deployments without requiring heavy frameworks like PyTorch or TensorFlow in production.
π» Quickstart
π¦ Embedded Metadata (sep_meta)
Every .onnx file in this repository is self-contained β no separate sidecar .json file is needed. All inference parameters are embedded directly inside the ONNX model as a metadata_props entry with key sep_meta.
Reading metadata with Python (onnxruntime)
import json
import onnxruntime as ort
sess = ort.InferenceSession("Reverb_HQ_By_FoxJoy.onnx", providers=["CPUExecutionProvider"])
meta_map = sess.get_modelmeta().custom_metadata_map # dict[str, str]
sep_meta = json.loads(meta_map["sep_meta"])
print(sep_meta["arch"]) # "MDX" | "VR" | "ROFORMER"
print(sep_meta["primary_stem"]) # "No Reverb"
print(sep_meta["secondary_stem"]) # "Reverb"
Reading metadata with Python (onnx library, no ORT session)
import json
import onnx
model = onnx.load("UVR-MDX-NET-Inst_HQ_1.onnx")
meta_map = {p.key: p.value for p in model.metadata_props}
sep_meta = json.loads(meta_map["sep_meta"])
Reading metadata with C (fast byte-scan, no ORT dependency)
The sep_meta key is stored near the end of the ONNX protobuf binary, after all weight initializers. You can extract it with a simple byte scan β no need to load the entire model into memory:
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
/* Returns a heap-allocated JSON string, or NULL on failure.
Caller must free() the result. */
char* read_onnx_sep_meta(const char* onnx_path) {
FILE* f = fopen(onnx_path, "rb");
if (!f) return NULL;
/* Read the last 64 KB β sep_meta is always at the end */
const size_t SCAN_SIZE = 64 * 1024;
fseek(f, 0, SEEK_END);
long file_size = ftell(f);
size_t read_size = (file_size < (long)SCAN_SIZE) ? (size_t)file_size : SCAN_SIZE;
long offset = file_size - (long)read_size;
fseek(f, offset, SEEK_SET);
char* buf = (char*)malloc(read_size + 1);
if (!buf) { fclose(f); return NULL; }
size_t n = fread(buf, 1, read_size, f);
fclose(f);
buf[n] = '\0';
/* Find the marker string "sep_meta" */
const char* marker = "sep_meta";
char* pos = buf;
char* found = NULL;
while ((pos = (char*)memchr(pos, 's', buf + n - pos)) != NULL) {
if ((size_t)(buf + n - pos) >= strlen(marker) &&
memcmp(pos, marker, strlen(marker)) == 0) {
found = pos; /* keep the last occurrence */
}
pos++;
}
if (!found) { free(buf); return NULL; }
/* After "sep_meta" the protobuf stores the value string.
Skip past the key and any protobuf length bytes until '{' */
char* json_start = found + strlen(marker);
while (json_start < buf + n && *json_start != '{') json_start++;
if (json_start >= buf + n) { free(buf); return NULL; }
/* Find matching closing '}' */
int depth = 0;
char* p = json_start;
while (p < buf + n) {
if (*p == '{') depth++;
else if (*p == '}') { depth--; if (depth == 0) break; }
p++;
}
if (depth != 0) { free(buf); return NULL; }
size_t json_len = (size_t)(p - json_start) + 1;
char* result = (char*)malloc(json_len + 1);
memcpy(result, json_start, json_len);
result[json_len] = '\0';
free(buf);
return result;
}
/* Example usage */
int main(void) {
char* meta = read_onnx_sep_meta("Reverb_HQ_By_FoxJoy.onnx");
if (meta) {
printf("%s\n", meta);
free(meta);
}
return 0;
}
π¬ sep_meta JSON Schema
The sep_meta value is a compact JSON object. Fields differ by architecture:
MDX-Net models
{
"arch": "MDX",
"primary_stem": "Vocals",
"secondary_stem": "Instrumental",
"sample_rate": 44100,
"n_fft": 7680,
"hop_length": 1024,
"dim_f": 3072,
"dim_t": 256,
"compensate": 1.021,
"overlap": 0.25
}
| Field | Type | Description |
|---|---|---|
n_fft |
int | STFT window size |
hop_length |
int | STFT hop length |
dim_f |
int | Frequency bins fed to the network |
dim_t |
int | Time frames per chunk |
compensate |
float | Amplitude gain applied after iSTFT |
overlap |
float | Overlap ratio between consecutive chunks (0β0.75) |
ONNX I/O:
- Input:
inputβ shape(1, 4, dim_f, dim_t)β four channels:[real_L, imag_L, real_R, imag_R] - Output:
outputβ same shape as input
VR Architecture models
{
"arch": "VR",
"primary_stem": "No Reverb",
"secondary_stem": "Reverb",
"sample_rate": 44100,
"vr_model_param": "4band_v3",
"bins": 672,
"window_size": 512,
"is_vr51": true,
"nn_arch_size": 218,
"model_capacity": [32, 128],
"band_params": {
"1": {"sr": 11025, "n_fft": 2048, "trim": 0, "window": "hann", "agg": 10},
"2": {"sr": 22050, "n_fft": 2048, "trim": 48, "window": "hann", "agg": 10},
"3": {"sr": 44100, "n_fft": 2048, "trim": 48, "window": "hann", "agg": 10},
"4": {"sr": 44100, "n_fft": 8192, "trim": 0, "window": "hann", "agg": 10}
}
}
| Field | Type | Description |
|---|---|---|
vr_model_param |
string | Band config preset name |
bins |
int | Total frequency bins after multi-band STFT |
window_size |
int | Window size for the primary STFT pass |
is_vr51 |
bool | true = CascadedNet (UVR v5.1), false = v4 |
nn_arch_size |
int | Network depth parameter |
band_params |
object | Per-band STFT config used to build the input spectrogram |
ONNX I/O:
- Input:
inputβ shape(batch, 2, bins, frames)β stereo multi-band magnitude spectrogram - Output:
outputβ shape(batch, 2, bins, frames)β predicted source mask
Roformer models (MelBandRoformer / BSRoformer)
Note: Roformer ONNX exports contain only the transformer core. STFT and iSTFT are not part of the ONNX graph (complex ops are not exportable). The inference wrapper must perform STFT β band-split β ONNX β scatter β iSTFT around the model.
{
"arch": "ROFORMER",
"roformer_type": "MelBandRoformer",
"primary_stem": "No Reverb",
"secondary_stem": "Reverb",
"sample_rate": 44100,
"chunk_size": 264600,
"hop_length": 1024,
"n_fft": 2048,
"dim_t": 256,
"num_bands": 64,
"dim": 128,
"frames": 256,
"n_freq_indices": 384,
"num_bins": 1025,
"n_full_freqs": 1025,
"overlap": 2,
"freq_indices": [...],
"num_bands_per_freq": [...]
}
| Field | Type | Description |
|---|---|---|
chunk_size |
int | Audio samples per inference chunk |
hop_length / n_fft |
int | STFT params applied outside the ONNX graph |
num_bands |
int | Number of mel bands |
frames |
int | Time frames (= dim_t) |
n_freq_indices |
int | Number of selected frequency bins after mel gather |
freq_indices |
int[] | Frequency bin indices to gather from full STFT before passing to network |
num_bands_per_freq |
int[] | Band assignment per frequency bin (for scatter-add in iSTFT) |
overlap |
int | Number of overlapping inference passes per chunk |
ONNX I/O:
- Input:
x_bandsβ shape(1, frames, num_bands, dim)β band-split STFT features - Output:
masksβ shape(1, 1, n_freq_indices, frames, 2)β complex mask
π Quickstart with audio-separator
The recommended way to run these models is audio-separator β the same inference engine that powers Ultimate Vocal Remover.
All .onnx and .pth files in this repo are fully compatible with audio-separator out of the box.
Installation
# CPU only (macOS / Linux / Windows)
pip install "audio-separator[cpu]"
# Nvidia GPU (CUDA)
pip install "audio-separator[gpu]"
# Apple Silicon (CoreML acceleration)
pip install "audio-separator[cpu]" # CoreML is auto-detected on macOS
CLI usage
Download the model file from the Files and versions tab, then point --model_file_dir at the folder containing it:
# Remove reverb from a mix
audio-separator mix.wav \
--model_filename Reverb_HQ_By_FoxJoy.onnx \
--model_file_dir /path/to/downloaded/models \
--output_dir ./output \
--output_format WAV
# Vocal / instrumental separation
audio-separator mix.wav \
--model_filename UVR-MDX-NET-Inst_HQ_3.onnx \
--model_file_dir /path/to/downloaded/models \
--output_dir ./output
Python API
from audio_separator.separator import Separator
separator = Separator(
model_file_dir="/path/to/downloaded/models", # folder with the .onnx files
output_dir="./output",
output_format="WAV",
)
# Load and run β arch is auto-detected from the file extension
separator.load_model("Reverb_HQ_By_FoxJoy.onnx")
output_files = separator.separate("mix.wav")
print("Output:", output_files)
Adjusting inference parameters
audio-separator exposes the most important per-arch parameters:
# MDX β tune overlap and segment size
separator = Separator(
model_file_dir="/path/to/downloaded/models",
mdx_params={
"hop_length": 1024,
"segment_size": 256, # dim_t from sep_meta
"overlap": 0.25, # matches sep_meta default
"batch_size": 1,
"enable_denoise": False,
}
)
# VR β tune aggression and window size
separator = Separator(
model_file_dir="/path/to/downloaded/models",
vr_params={
"batch_size": 1,
"window_size": 512, # 320 = slower but better
"aggression": 5, # 0-100, higher = more aggressive separation
"enable_tta": False,
"high_end_process": False,
}
)
separator.load_model("UVR-DeEcho-DeReverb.onnx")
output_files = separator.separate("mix.wav")
Note on
sep_meta:audio-separatorreads model parameters from its own internalmdx_model_data.jsonlookup table β it does not read the embeddedsep_metametadata. Thesep_metafield is provided for custom inference pipelines and other runtimes (see sections below). embed_sep_meta(onnx_path: str, meta: dict) -> None:
model = onnx.load(onnx_path)
# Remove stale entry if present
existing = [p for p in model.metadata_props if p.key != "sep_meta"]
del model.metadata_props[:]
model.metadata_props.extend(existing)
# Add new entry
entry = model.metadata_props.add()
entry.key = "sep_meta"
entry.value = json.dumps(meta, separators=(",", ":"))
onnx.save_model(model, onnx_path)
Example β MDX model
embed_sep_meta("my_model.onnx", {
"arch": "MDX",
"primary_stem": "Vocals",
"secondary_stem": "Instrumental",
"sample_rate": 44100,
"n_fft": 7680,
"hop_length": 1024,
"dim_f": 3072,
"dim_t": 256,
"compensate": 1.021,
"overlap": 0.25,
})
π Requirements
onnxruntime >= 1.16
numpy >= 1.24
soundfile >= 0.12
For metadata-only access (no inference):
onnx >= 1.14
π License
The repository structure and conversion scripts are licensed under the MIT License. Please check the original model licenses before using them for commercial purposes.
π Acknowledgments & Credits
Special thanks to the authors and maintainers of: