open-rvq-encoder-minimax-music3-169m β€” soft-distilled + calibrated (hotstep-v1)

Mothersuperior/open-rvq-encoder-minimax-music3-169m-pooled-v4 further-trained with a pure soft-distillation objective (soft cross-entropy against the teacher top-50 distributions shipped in the 8k corpus) plus learnable per-head logit temperatures, then selected on a real-audio gate that no other release in this lineage measures. Architecture, config, and file format are identical to the base β€” this is a drop-in replacement.

Three things in this release are useful even if you never load the weights:

1. Corpus consumption fixes (apply these before training on the 8k corpus)

uint16 wrap in probs.npz semantic ids. The semantic head's idx values were stored as raw LM token ids (code + 151675) in a uint16, which silently wraps them mod 65536. They arrive as code + 20603. Subtract 20603 and mask anything outside [0, 16384) (the EOS token wraps to βˆ’5). Verified: after unwrapping, the sampled ground-truth code lands inside the teacher top-50 for ~100% of frames; consumed naively, the ids are garbage and soft-distillation losses are meaningless.

Stitched-timeline latent alignment. The corpus audio is stitched from 200-frame DiT windows at a 100-frame / 345-latent hop (documented in the corpus card), while the naive uniform mapping (latent = frame * 441 // 128 β‰ˆ 344.53/100 frames) drifts ~1 latent per 800 frames. Measured effect: semantic top-1 against ground-truth codes decays from 0.41 at track start to 0.06 by frame 2400 under the uniform rule. The correct frameβ†’latent mapping for this corpus is:

def stitched_lat(i):                      # frame index -> DAV latent index
    k = 0 if i < 125 else (i - 25) // 100 # chunk ownership (from local frame 25)
    return 345 * k + ((i - 100 * k) * 441) // 128 + 1

The +1 is a constant global offset, verified empirically as best at every probed position on every probed record (top-1 flat at 0.4–0.6 across track positions once applied). The corpus's own DAV latents are exact β€” recomputing them from the shipped FLAC through the official dav.pth reproduces them at cosine 1.0 β€” the drift lives in the audio-vs-codes timeline of chunked rendering. Real audio has no such drift: the uniform mapping remains correct at inference.

2. Calibrated logits (folded)

The released v4/pooled-v4 readout logits are muP-sharp and uncalibrated: harmless for argmax/top-K use (what the reference adapter and the ComfyUI/diffusers nodes do), but cross-entropy and softmax probabilities computed from them are meaningless (semantic soft-CE ~59 nats at temperature 1; per-frame logit std ~45). This release trained learnable per-head temperatures jointly with the model (initialised from a line search: ~16 semantic, ~2 acoustic) and folded them into the head weights at export. Argmax and top-K ordering are unchanged; softmax(logits) and CE are now meaningful out of the box. The final folded values are in calibration.json.

3. The weights

Training: pure soft-CE vs teacher top-50 (all 8 heads; depth decoder teacher-forced), 20,000 steps at effective batch 64 (16 Γ— 4 grad-accum), lr 1e-4 cosine with 200-step warmup, bf16, single RTX 5090, on the drift-corrected + unwrapped 8k corpus (records < 8 s filtered; 7,726 train / 394 holdout by salted-hash split). Trainer included in training/.

In-domain holdout (8k corpus, drift-corrected loader, calibrated CE)

holdout metric pooled-v4 (start) this release (step 18500) change
semantic soft-CE (vs teacher) 3.299 3.039 βˆ’7.9%
acoustic soft-CE 4.618 4.413 βˆ’4.4%
semantic hard-CE 2.336 2.154 βˆ’7.8%
acoustic hard-CE 3.925 3.763 βˆ’4.1%
semantic top-1 41.9% 42.2% +0.3 pt
acoustic top-1 16.8% 17.6% +0.8 pt

In-domain argmax barely moves β€” the interesting result is below.

Real-audio gate (13 commercial studio recordings, frozen official LM/depth decoder base-CE)

The metric nobody else in the lineage reports: export codes from real recordings, measure the frozen official model's cross-entropy on them. Lower = the encoder's codes look more like codes the official model would emit. All prior rows reproduced on the same harness:

encoder params sem CE sem t1 ac CE ac t1
Serveurperso community 41M 41M 8.447 5.8% 6.763 2.3%
Mothersuperior pooled-v3 41M 7.261 8.5% 6.347 4.5%
SimpleTuner v4 169M 7.158 8.9% 5.638 5.8%
Mothersuperior pooled-v4 169M 6.838 10.9% 5.373 7.5%
this release 169M 5.866 16.6% 4.725 10.9%

The largest single-step improvement in the lineage on this metric β€” and it happens at the argmax level (temperature folding cannot cause it), despite near-flat in-domain argmax accuracy. Soft distillation appears to act as regularisation against synthetic-audio overfit, transferring to real recordings. Caveat: semantic unique-code usage on the real-audio set drops ~14% vs the base (6,490 vs 7,515 over 64k frames; acoustic unchanged) β€” mild concentration, far from collapse, but stated for transparency.

Ear-verified on frame-matched replay renders of five real commercial recordings across four genres (indie rock, pop, alt-metal, electronic): an audible improvement over pooled-v4 on all of them, though subtler to the ear than the pooled-v4 β†’ v4 step was β€” consistent with diminishing returns from a synthetic-only corpus at this size. Best in the lineage by both metric and ear; the remaining real-audio fidelity gap likely needs real-audio objectives (consistency training, LM-prior regularisation, reconstruction through the frozen generator) rather than more synthetic epochs.

Usage

Identical to the base model. Load with SimpleTuner's minimax_music3_reference_adapter.py:

from minimax_music3_reference_adapter import MiniMaxMusic3ReferenceAdapter
adapter = MiniMaxMusic3ReferenceAdapter.from_files(
    "rvq_encoder.safetensors", "rvq_encoder_config.json", "dav.pth")
codes = adapter.predict_codes(waveform, sample_rate)   # [frames, 8]

dav.pth is the official DAV encoder from MiniMaxAI/MiniMax-Music3. Strict load_state_dict β€” the checkpoint keys match the base exactly.

Further training

training/rvq_distill_train.py is the self-contained trainer used for this release: 8k-corpus zip layout, both consumption fixes above built in, warm start from any encoder in this lineage, pure-soft or mixed soft/hard loss, learnable temperatures with fold-at-save, holdout eval + best-checkpoint selection. Depends only on torch/numpy/safetensors plus the reference adapter module for the model classes.

Credits

This encoder exists because of a chain of community work, each stage building on the last:

  • MiniMax β€” MiniMax Music 3 and the DAV encoder. All weights and training data derive from it; use is subject to its terms.
  • Serveurperso β€” the original community proof: a 41M single-GPU encoder demonstrating that exact token agreement was not required, and the first independent derivation of the stitched-timeline alignment contract.
  • bghira / SimpleTuner β€” the encoder architecture (v1–v4, causal depth decoder), the reference adapter, the trainer, and the reverse-distillation trace corpus.
  • Mothersuperior β€” the 8k distillation corpus with teacher top-50 distributions (which make this release's objective possible), and the pooled-v3/pooled-v4 fine-tunes this release warm-starts from.
  • HOT-Step CPP β€” this release: the soft-distillation training run, the corpus consumption fixes, logit calibration, and the real-audio gate.

Use is subject to the MiniMax Music 3 model terms and the source datasets' terms.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for scragnog/open-rvq-encoder-minimax-music3-169m-hotstep-v1

Dataset used to train scragnog/open-rvq-encoder-minimax-music3-169m-hotstep-v1