Buckets:
Recovered audio-loss trainer: joint_v6.py
This is the original research script recovered on 2026-09-19 from
/workspace/tok/full/joint_v6.py. It is the trainer used for the v6–v9 audio-loss
weight sweep, not the later reference-to-song adapter trainer. The file is
byte-for-byte preserved; no training logic or paths were changed for this release.
Original SHA-256: 07a53c28262cfcda90ee1a0d37facc82ad6a4de4d5c448305f4ae69202836ca9.
Validation for this recovery was source-hash equality, Python syntax checking, and lookup-table shape/range checks. The original GPU experiment was not rerun as part of publication; this is not a newly validated portable training package.
Objective and sweep
The script jointly trains the MERT-feature tokenizer head and the acoustic decoder LoRA. Real-audio flow loss backpropagates through straight-through token embeddings. A frozen VAE decoder adds log-mel, multi-resolution STFT, and stereo width losses on an estimated clean latent crop. Minted examples supply head soft-label CE and 25% of decoder flow windows.
The same script runs every audio-loss variant:
| Run name | AUX_W |
|---|---|
| joint_v6 | 0.3 |
| joint_v7 | 1.0 |
| joint_v8 | 2.0 |
| joint_v9 | 4.0 |
AUX_W defaults to zero, so set it explicitly to enable the waveform loss.
Other defaults: AUX_TMAX=0.4, AUX_FR=150, AUX_MARGIN=25, 512-frame training
windows, and 25 frames/second. The waveform term applies to real windows only
when sampled flow time is at most AUX_TMAX. It can also be skipped when the
original audio cannot be resolved or the requested crop is too short.
Original environment and required inputs
Use the older yue2-infer runtime described in the repository README (commit
92a73cc7, original Python 3.12 / torch 2.10 / CUDA 12.8 environment), plus numpy,
soundfile, torchaudio, and ffmpeg. Current YuE2 APIs may differ. Do not replace a
working modern training environment with these older dependencies.
The original script has hard-coded paths. Recreate them or edit a working copy:
/workspace/hf/hub/models--m-a-p--YuE2-3B/snapshots/...and the correspondingYuE2-Vaesnapshot. Both must already exist. The script takes the first glob match; isolate one intended snapshot to avoid accidental version selection./workspace/tok/full/feats/<minted_id>.npy: minted MERT L20 features, either[T,1024]or the legacy[4,T,1024]layout (index 3 is selected)./workspace/yue2-corpus/tracks/<minted_id>/:semantic.npy,latent.npy, andrequest.jsonwithstyle,lyrics, andseed. The small AR regularizer pack alone is insufficient for this trainer: full features and latents are needed.- Copy
assets/sem_nbr_idx.npyandassets/sem_nbr_cos.npyfrom this repository to/workspace/tok/full/. These are semantic-token neighbor lookup tables, not song data. Their exact hashes are injoint_v6_release.json. RPpoints to prepared real recordings. Each child directory needsmert.npy[T,1024],lat.npy[T,64], and integerprefix.npy. Training songs require at least 512 frames. SetHOLDto existing prepared directory names, separated by commas; supply at least one sufficiently long held-out recording.- Edit
audio_path()to map each prepared name to its exact original audio. The historical mapping recognizescoheed__<name>and other<tag>__<name>layouts. Without a working mapping the auxiliary audio loss may silently disappear. Check nonzeroauxvalues on eligible real updates; legitimate zero values also occur for minted windows and larger flow times. - Create
/workspace/tok/full/listen_real/for final renders. Run names should be unique: this historical script does not implement safe optimizer resume.
Checkpoint format and example invocation
Unlike the other adapted scripts in this repository, the recovered original
calls torch.load and expects .pt dictionaries. For a safetensors release,
convert a trusted checkpoint with the existing ckpt_io.py helper first:
import torch
from ckpt_io import load_ckpt
torch.save(load_ckpt("tokenizer_head_v5_30k.safetensors"), "head_init.pt")
The historical sweep started from the pretrained head corresponding to
tokenizer_head_v5_30k.safetensors and nar_lora_joint_v4.pt. All variants used
the same initial weights, rather than chaining v6 into v7 into v8 into v9.
For an already prepared custom dataset, after resolving the paths above:
mkdir -p /workspace/tok/full/listen_real
RP=/path/to/prepared_real HOLD=heldout_song_1,heldout_song_2 \
MINTED_CAP=4000 AUX_W=4.0 AUX_TMAX=0.4 AUX_FR=150 \
python scripts/joint_v6.py my_audio_loss_run 3000 1 1 \
/path/to/head_init.pt /path/to/nar_lora_joint_v4.pt 32
This example is a historical recipe, not a claim that 3,000 steps or those
hyperparameters are optimal for a new dataset. best is selected by held-out
latent flow loss, not the auxiliary spectral metric or a listening score.
The outputs include head/LoRA best and last weights and held-out reconstructions;
they are not new-song artist-conditioning demonstrations.
Xet Storage Details
- Size:
- 5.13 kB
- Xet hash:
- 625627ed7ff636b61214fc03f567cc7e983996d03b5979faf0afaa6a1cc5767f
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.