AudioSeparatorONNX / README.md
raydotac's picture
Update README.md
c02cba6 verified
|
Raw History Blame
12.7 kB
metadata
license: mit
pipeline_tag: audio-to-audio
tags:
  - audio-separation
  - vocal-remover
  - stem-separation
  - onnx
  - onnxruntime

AudioSeparatorONNX

Collection of popular audio separation and vocal removal models converted to ONNX format, optimized for high-performance CPU/GPU inference using onnxruntime.

πŸ“Œ Overview

This repository provides pre-converted ONNX versions of popular music source separation (MSS) and vocal removal models (such as UVR5 / Ultimate Vocal Remover models, MDX-Net, VR Architecture, Demucs, etc.).

Using ONNX models allows for:

  • Cross-platform compatibility (Python, C++, C#, Rust, JS, etc.).
  • Faster execution with onnxruntime using CPU, CUDA, DirectML, CoreML, or TensorRT.
  • Lightweight deployments without requiring heavy frameworks like PyTorch or TensorFlow in production.

πŸ’» Quickstart

πŸ“¦ Embedded Metadata (sep_meta)

Every .onnx file in this repository is self-contained β€” no separate sidecar .json file is needed. All inference parameters are embedded directly inside the ONNX model as a metadata_props entry with key sep_meta.

Reading metadata with Python (onnxruntime)

import json
import onnxruntime as ort

sess = ort.InferenceSession("Reverb_HQ_By_FoxJoy.onnx", providers=["CPUExecutionProvider"])
meta_map = sess.get_modelmeta().custom_metadata_map   # dict[str, str]
sep_meta = json.loads(meta_map["sep_meta"])

print(sep_meta["arch"])           # "MDX" | "VR" | "ROFORMER"
print(sep_meta["primary_stem"])   # "No Reverb"
print(sep_meta["secondary_stem"]) # "Reverb"

Reading metadata with Python (onnx library, no ORT session)

import json
import onnx

model = onnx.load("UVR-MDX-NET-Inst_HQ_1.onnx")
meta_map = {p.key: p.value for p in model.metadata_props}
sep_meta = json.loads(meta_map["sep_meta"])

Reading metadata with C (fast byte-scan, no ORT dependency)

The sep_meta key is stored near the end of the ONNX protobuf binary, after all weight initializers. You can extract it with a simple byte scan β€” no need to load the entire model into memory:

#include <stdio.h>
#include <stdlib.h>
#include <string.h>

/* Returns a heap-allocated JSON string, or NULL on failure.
   Caller must free() the result. */
char* read_onnx_sep_meta(const char* onnx_path) {
    FILE* f = fopen(onnx_path, "rb");
    if (!f) return NULL;

    /* Read the last 64 KB β€” sep_meta is always at the end */
    const size_t SCAN_SIZE = 64 * 1024;
    fseek(f, 0, SEEK_END);
    long file_size = ftell(f);
    size_t read_size = (file_size < (long)SCAN_SIZE) ? (size_t)file_size : SCAN_SIZE;
    long offset = file_size - (long)read_size;
    fseek(f, offset, SEEK_SET);

    char* buf = (char*)malloc(read_size + 1);
    if (!buf) { fclose(f); return NULL; }
    size_t n = fread(buf, 1, read_size, f);
    fclose(f);
    buf[n] = '\0';

    /* Find the marker string "sep_meta" */
    const char* marker = "sep_meta";
    char* pos = buf;
    char* found = NULL;
    while ((pos = (char*)memchr(pos, 's', buf + n - pos)) != NULL) {
        if ((size_t)(buf + n - pos) >= strlen(marker) &&
            memcmp(pos, marker, strlen(marker)) == 0) {
            found = pos;  /* keep the last occurrence */
        }
        pos++;
    }
    if (!found) { free(buf); return NULL; }

    /* After "sep_meta" the protobuf stores the value string.
       Skip past the key and any protobuf length bytes until '{' */
    char* json_start = found + strlen(marker);
    while (json_start < buf + n && *json_start != '{') json_start++;
    if (json_start >= buf + n) { free(buf); return NULL; }

    /* Find matching closing '}' */
    int depth = 0;
    char* p = json_start;
    while (p < buf + n) {
        if (*p == '{') depth++;
        else if (*p == '}') { depth--; if (depth == 0) break; }
        p++;
    }
    if (depth != 0) { free(buf); return NULL; }

    size_t json_len = (size_t)(p - json_start) + 1;
    char* result = (char*)malloc(json_len + 1);
    memcpy(result, json_start, json_len);
    result[json_len] = '\0';

    free(buf);
    return result;
}

/* Example usage */
int main(void) {
    char* meta = read_onnx_sep_meta("Reverb_HQ_By_FoxJoy.onnx");
    if (meta) {
        printf("%s\n", meta);
        free(meta);
    }
    return 0;
}

πŸ”¬ sep_meta JSON Schema

The sep_meta value is a compact JSON object. Fields differ by architecture:

MDX-Net models

{
  "arch":           "MDX",
  "primary_stem":   "Vocals",
  "secondary_stem": "Instrumental",
  "sample_rate":    44100,
  "n_fft":          7680,
  "hop_length":     1024,
  "dim_f":          3072,
  "dim_t":          256,
  "compensate":     1.021,
  "overlap":        0.25
}
Field Type Description
n_fft int STFT window size
hop_length int STFT hop length
dim_f int Frequency bins fed to the network
dim_t int Time frames per chunk
compensate float Amplitude gain applied after iSTFT
overlap float Overlap ratio between consecutive chunks (0–0.75)

ONNX I/O:

  • Input: input β€” shape (1, 4, dim_f, dim_t) β€” four channels: [real_L, imag_L, real_R, imag_R]
  • Output: output β€” same shape as input

VR Architecture models

{
  "arch":           "VR",
  "primary_stem":   "No Reverb",
  "secondary_stem": "Reverb",
  "sample_rate":    44100,
  "vr_model_param": "4band_v3",
  "bins":           672,
  "window_size":    512,
  "is_vr51":        true,
  "nn_arch_size":   218,
  "model_capacity": [32, 128],
  "band_params": {
    "1": {"sr": 11025, "n_fft": 2048, "trim": 0, "window": "hann", "agg": 10},
    "2": {"sr": 22050, "n_fft": 2048, "trim": 48, "window": "hann", "agg": 10},
    "3": {"sr": 44100, "n_fft": 2048, "trim": 48, "window": "hann", "agg": 10},
    "4": {"sr": 44100, "n_fft": 8192, "trim": 0, "window": "hann", "agg": 10}
  }
}
Field Type Description
vr_model_param string Band config preset name
bins int Total frequency bins after multi-band STFT
window_size int Window size for the primary STFT pass
is_vr51 bool true = CascadedNet (UVR v5.1), false = v4
nn_arch_size int Network depth parameter
band_params object Per-band STFT config used to build the input spectrogram

ONNX I/O:

  • Input: input β€” shape (batch, 2, bins, frames) β€” stereo multi-band magnitude spectrogram
  • Output: output β€” shape (batch, 2, bins, frames) β€” predicted source mask

Roformer models (MelBandRoformer / BSRoformer)

Note: Roformer ONNX exports contain only the transformer core. STFT and iSTFT are not part of the ONNX graph (complex ops are not exportable). The inference wrapper must perform STFT β†’ band-split β†’ ONNX β†’ scatter β†’ iSTFT around the model.

{
  "arch":               "ROFORMER",
  "roformer_type":      "MelBandRoformer",
  "primary_stem":       "No Reverb",
  "secondary_stem":     "Reverb",
  "sample_rate":        44100,
  "chunk_size":         264600,
  "hop_length":         1024,
  "n_fft":              2048,
  "dim_t":              256,
  "num_bands":          64,
  "dim":                128,
  "frames":             256,
  "n_freq_indices":     384,
  "num_bins":           1025,
  "n_full_freqs":       1025,
  "overlap":            2,
  "freq_indices":       [...],
  "num_bands_per_freq": [...]
}
Field Type Description
chunk_size int Audio samples per inference chunk
hop_length / n_fft int STFT params applied outside the ONNX graph
num_bands int Number of mel bands
frames int Time frames (= dim_t)
n_freq_indices int Number of selected frequency bins after mel gather
freq_indices int[] Frequency bin indices to gather from full STFT before passing to network
num_bands_per_freq int[] Band assignment per frequency bin (for scatter-add in iSTFT)
overlap int Number of overlapping inference passes per chunk

ONNX I/O:

  • Input: x_bands β€” shape (1, frames, num_bands, dim) β€” band-split STFT features
  • Output: masks β€” shape (1, 1, n_freq_indices, frames, 2) β€” complex mask

πŸš€ Quickstart with audio-separator

The recommended way to run these models is audio-separator β€” the same inference engine that powers Ultimate Vocal Remover.

All .onnx and .pth files in this repo are fully compatible with audio-separator out of the box.

Installation

# CPU only (macOS / Linux / Windows)
pip install "audio-separator[cpu]"

# Nvidia GPU (CUDA)
pip install "audio-separator[gpu]"

# Apple Silicon (CoreML acceleration)
pip install "audio-separator[cpu]"   # CoreML is auto-detected on macOS

CLI usage

Download the model file from the Files and versions tab, then point --model_file_dir at the folder containing it:

# Remove reverb from a mix
audio-separator mix.wav \
  --model_filename Reverb_HQ_By_FoxJoy.onnx \
  --model_file_dir /path/to/downloaded/models \
  --output_dir ./output \
  --output_format WAV

# Vocal / instrumental separation
audio-separator mix.wav \
  --model_filename UVR-MDX-NET-Inst_HQ_3.onnx \
  --model_file_dir /path/to/downloaded/models \
  --output_dir ./output

Python API

from audio_separator.separator import Separator

separator = Separator(
    model_file_dir="/path/to/downloaded/models",  # folder with the .onnx files
    output_dir="./output",
    output_format="WAV",
)

# Load and run β€” arch is auto-detected from the file extension
separator.load_model("Reverb_HQ_By_FoxJoy.onnx")
output_files = separator.separate("mix.wav")
print("Output:", output_files)

Adjusting inference parameters

audio-separator exposes the most important per-arch parameters:

# MDX β€” tune overlap and segment size
separator = Separator(
    model_file_dir="/path/to/downloaded/models",
    mdx_params={
        "hop_length":    1024,
        "segment_size":  256,    # dim_t from sep_meta
        "overlap":       0.25,   # matches sep_meta default
        "batch_size":    1,
        "enable_denoise": False,
    }
)

# VR β€” tune aggression and window size
separator = Separator(
    model_file_dir="/path/to/downloaded/models",
    vr_params={
        "batch_size":    1,
        "window_size":   512,    # 320 = slower but better
        "aggression":    5,      # 0-100, higher = more aggressive separation
        "enable_tta":    False,
        "high_end_process": False,
    }
)

separator.load_model("UVR-DeEcho-DeReverb.onnx")
output_files = separator.separate("mix.wav")

Note on sep_meta: audio-separator reads model parameters from its own internal mdx_model_data.json lookup table β€” it does not read the embedded sep_meta metadata. The sep_meta field is provided for custom inference pipelines and other runtimes (see sections below). embed_sep_meta(onnx_path: str, meta: dict) -> None:

    model = onnx.load(onnx_path)
    # Remove stale entry if present
    existing = [p for p in model.metadata_props if p.key != "sep_meta"]
    del model.metadata_props[:]
    model.metadata_props.extend(existing)
    # Add new entry
    entry = model.metadata_props.add()
    entry.key   = "sep_meta"
    entry.value = json.dumps(meta, separators=(",", ":"))
    onnx.save_model(model, onnx_path)

Example β€” MDX model

embed_sep_meta("my_model.onnx", {
    "arch":           "MDX",
    "primary_stem":   "Vocals",
    "secondary_stem": "Instrumental",
    "sample_rate":    44100,
    "n_fft":          7680,
    "hop_length":     1024,
    "dim_f":          3072,
    "dim_t":          256,
    "compensate":     1.021,
    "overlap":        0.25,
})

πŸ“‹ Requirements

onnxruntime >= 1.16
numpy >= 1.24
soundfile >= 0.12

For metadata-only access (no inference):

onnx >= 1.14

πŸ“„ License

The repository structure and conversion scripts are licensed under the MIT License. Please check the original model licenses before using them for commercial purposes.


πŸ™ Acknowledgments & Credits

Special thanks to the authors and maintainers of: