osomwolf/SlopPrompt-bucket / scripts /joint_v6_README.md
osomwolf's picture
|
download
raw
5.13 kB

Recovered audio-loss trainer: joint_v6.py

This is the original research script recovered on 2026-09-19 from /workspace/tok/full/joint_v6.py. It is the trainer used for the v6–v9 audio-loss weight sweep, not the later reference-to-song adapter trainer. The file is byte-for-byte preserved; no training logic or paths were changed for this release. Original SHA-256: 07a53c28262cfcda90ee1a0d37facc82ad6a4de4d5c448305f4ae69202836ca9.

Validation for this recovery was source-hash equality, Python syntax checking, and lookup-table shape/range checks. The original GPU experiment was not rerun as part of publication; this is not a newly validated portable training package.

Objective and sweep

The script jointly trains the MERT-feature tokenizer head and the acoustic decoder LoRA. Real-audio flow loss backpropagates through straight-through token embeddings. A frozen VAE decoder adds log-mel, multi-resolution STFT, and stereo width losses on an estimated clean latent crop. Minted examples supply head soft-label CE and 25% of decoder flow windows.

The same script runs every audio-loss variant:

Run name AUX_W
joint_v6 0.3
joint_v7 1.0
joint_v8 2.0
joint_v9 4.0

AUX_W defaults to zero, so set it explicitly to enable the waveform loss. Other defaults: AUX_TMAX=0.4, AUX_FR=150, AUX_MARGIN=25, 512-frame training windows, and 25 frames/second. The waveform term applies to real windows only when sampled flow time is at most AUX_TMAX. It can also be skipped when the original audio cannot be resolved or the requested crop is too short.

Original environment and required inputs

Use the older yue2-infer runtime described in the repository README (commit 92a73cc7, original Python 3.12 / torch 2.10 / CUDA 12.8 environment), plus numpy, soundfile, torchaudio, and ffmpeg. Current YuE2 APIs may differ. Do not replace a working modern training environment with these older dependencies.

The original script has hard-coded paths. Recreate them or edit a working copy:

  • /workspace/hf/hub/models--m-a-p--YuE2-3B/snapshots/... and the corresponding YuE2-Vae snapshot. Both must already exist. The script takes the first glob match; isolate one intended snapshot to avoid accidental version selection.
  • /workspace/tok/full/feats/<minted_id>.npy: minted MERT L20 features, either [T,1024] or the legacy [4,T,1024] layout (index 3 is selected).
  • /workspace/yue2-corpus/tracks/<minted_id>/: semantic.npy, latent.npy, and request.json with style, lyrics, and seed. The small AR regularizer pack alone is insufficient for this trainer: full features and latents are needed.
  • Copy assets/sem_nbr_idx.npy and assets/sem_nbr_cos.npy from this repository to /workspace/tok/full/. These are semantic-token neighbor lookup tables, not song data. Their exact hashes are in joint_v6_release.json.
  • RP points to prepared real recordings. Each child directory needs mert.npy [T,1024], lat.npy [T,64], and integer prefix.npy. Training songs require at least 512 frames. Set HOLD to existing prepared directory names, separated by commas; supply at least one sufficiently long held-out recording.
  • Edit audio_path() to map each prepared name to its exact original audio. The historical mapping recognizes coheed__<name> and other <tag>__<name> layouts. Without a working mapping the auxiliary audio loss may silently disappear. Check nonzero aux values on eligible real updates; legitimate zero values also occur for minted windows and larger flow times.
  • Create /workspace/tok/full/listen_real/ for final renders. Run names should be unique: this historical script does not implement safe optimizer resume.

Checkpoint format and example invocation

Unlike the other adapted scripts in this repository, the recovered original calls torch.load and expects .pt dictionaries. For a safetensors release, convert a trusted checkpoint with the existing ckpt_io.py helper first:

import torch
from ckpt_io import load_ckpt
torch.save(load_ckpt("tokenizer_head_v5_30k.safetensors"), "head_init.pt")

The historical sweep started from the pretrained head corresponding to tokenizer_head_v5_30k.safetensors and nar_lora_joint_v4.pt. All variants used the same initial weights, rather than chaining v6 into v7 into v8 into v9. For an already prepared custom dataset, after resolving the paths above:

mkdir -p /workspace/tok/full/listen_real
RP=/path/to/prepared_real HOLD=heldout_song_1,heldout_song_2 \
MINTED_CAP=4000 AUX_W=4.0 AUX_TMAX=0.4 AUX_FR=150 \
python scripts/joint_v6.py my_audio_loss_run 3000 1 1 \
  /path/to/head_init.pt /path/to/nar_lora_joint_v4.pt 32

This example is a historical recipe, not a claim that 3,000 steps or those hyperparameters are optimal for a new dataset. best is selected by held-out latent flow loss, not the auxiliary spectral metric or a listening score. The outputs include head/LoRA best and last weights and held-out reconstructions; they are not new-song artist-conditioning demonstrations.

Xet Storage Details

Size:
5.13 kB
·
Xet hash:
625627ed7ff636b61214fc03f567cc7e983996d03b5979faf0afaa6a1cc5767f

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.