Buckets:
| # Recovered audio-loss trainer: joint_v6.py | |
| This is the original research script recovered on 2026-09-19 from | |
| `/workspace/tok/full/joint_v6.py`. It is the trainer used for the v6–v9 audio-loss | |
| weight sweep, not the later reference-to-song adapter trainer. The file is | |
| byte-for-byte preserved; no training logic or paths were changed for this release. | |
| Original SHA-256: `07a53c28262cfcda90ee1a0d37facc82ad6a4de4d5c448305f4ae69202836ca9`. | |
| Validation for this recovery was source-hash equality, Python syntax checking, | |
| and lookup-table shape/range checks. The original GPU experiment was not rerun | |
| as part of publication; this is not a newly validated portable training package. | |
| ## Objective and sweep | |
| The script jointly trains the MERT-feature tokenizer head and the acoustic | |
| decoder LoRA. Real-audio flow loss backpropagates through straight-through token | |
| embeddings. A frozen VAE decoder adds log-mel, multi-resolution STFT, and stereo | |
| width losses on an estimated clean latent crop. Minted examples supply head | |
| soft-label CE and 25% of decoder flow windows. | |
| The same script runs every audio-loss variant: | |
| | Run name | `AUX_W` | | |
| | --- | ---: | | |
| | joint_v6 | 0.3 | | |
| | joint_v7 | 1.0 | | |
| | joint_v8 | 2.0 | | |
| | joint_v9 | 4.0 | | |
| **`AUX_W` defaults to zero**, so set it explicitly to enable the waveform loss. | |
| Other defaults: `AUX_TMAX=0.4`, `AUX_FR=150`, `AUX_MARGIN=25`, 512-frame training | |
| windows, and 25 frames/second. The waveform term applies to real windows only | |
| when sampled flow time is at most `AUX_TMAX`. It can also be skipped when the | |
| original audio cannot be resolved or the requested crop is too short. | |
| ## Original environment and required inputs | |
| Use the older `yue2-infer` runtime described in the repository README (commit | |
| `92a73cc7`, original Python 3.12 / torch 2.10 / CUDA 12.8 environment), plus numpy, | |
| soundfile, torchaudio, and ffmpeg. Current YuE2 APIs may differ. Do not replace a | |
| working modern training environment with these older dependencies. | |
| The original script has hard-coded paths. Recreate them or edit a working copy: | |
| - `/workspace/hf/hub/models--m-a-p--YuE2-3B/snapshots/...` and the corresponding | |
| `YuE2-Vae` snapshot. Both must already exist. The script takes the first glob | |
| match; isolate one intended snapshot to avoid accidental version selection. | |
| - `/workspace/tok/full/feats/<minted_id>.npy`: minted MERT L20 features, either | |
| `[T,1024]` or the legacy `[4,T,1024]` layout (index 3 is selected). | |
| - `/workspace/yue2-corpus/tracks/<minted_id>/`: `semantic.npy`, `latent.npy`, and | |
| `request.json` with `style`, `lyrics`, and `seed`. The small AR regularizer pack | |
| alone is insufficient for this trainer: full features and latents are needed. | |
| - Copy `assets/sem_nbr_idx.npy` and `assets/sem_nbr_cos.npy` from this repository | |
| to `/workspace/tok/full/`. These are semantic-token neighbor lookup tables, | |
| not song data. Their exact hashes are in `joint_v6_release.json`. | |
| - `RP` points to prepared real recordings. Each child directory needs `mert.npy` | |
| `[T,1024]`, `lat.npy` `[T,64]`, and integer `prefix.npy`. Training songs require | |
| at least 512 frames. Set `HOLD` to existing prepared directory names, separated | |
| by commas; supply at least one sufficiently long held-out recording. | |
| - Edit `audio_path()` to map each prepared name to its exact original audio. | |
| The historical mapping recognizes `coheed__<name>` and other `<tag>__<name>` | |
| layouts. Without a working mapping the auxiliary audio loss may silently | |
| disappear. Check nonzero `aux` values on eligible real updates; legitimate | |
| zero values also occur for minted windows and larger flow times. | |
| - Create `/workspace/tok/full/listen_real/` for final renders. Run names should | |
| be unique: this historical script does not implement safe optimizer resume. | |
| ## Checkpoint format and example invocation | |
| Unlike the other adapted scripts in this repository, the recovered original | |
| calls `torch.load` and expects `.pt` dictionaries. For a safetensors release, | |
| convert a trusted checkpoint with the existing `ckpt_io.py` helper first: | |
| ```python | |
| import torch | |
| from ckpt_io import load_ckpt | |
| torch.save(load_ckpt("tokenizer_head_v5_30k.safetensors"), "head_init.pt") | |
| ``` | |
| The historical sweep started from the pretrained head corresponding to | |
| `tokenizer_head_v5_30k.safetensors` and `nar_lora_joint_v4.pt`. All variants used | |
| the same initial weights, rather than chaining v6 into v7 into v8 into v9. | |
| For an already prepared custom dataset, after resolving the paths above: | |
| ```sh | |
| mkdir -p /workspace/tok/full/listen_real | |
| RP=/path/to/prepared_real HOLD=heldout_song_1,heldout_song_2 \ | |
| MINTED_CAP=4000 AUX_W=4.0 AUX_TMAX=0.4 AUX_FR=150 \ | |
| python scripts/joint_v6.py my_audio_loss_run 3000 1 1 \ | |
| /path/to/head_init.pt /path/to/nar_lora_joint_v4.pt 32 | |
| ``` | |
| This example is a historical recipe, not a claim that 3,000 steps or those | |
| hyperparameters are optimal for a new dataset. `best` is selected by held-out | |
| latent flow loss, not the auxiliary spectral metric or a listening score. | |
| The outputs include head/LoRA best and last weights and held-out reconstructions; | |
| they are not new-song artist-conditioning demonstrations. | |
Xet Storage Details
- Size:
- 5.13 kB
- Xet hash:
- 625627ed7ff636b61214fc03f567cc7e983996d03b5979faf0afaa6a1cc5767f
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.