# Recovered audio-loss trainer: joint_v6.py

This is the original research script recovered on 2026-09-19 from
`/workspace/tok/full/joint_v6.py`. It is the trainer used for the v6–v9 audio-loss
weight sweep, not the later reference-to-song adapter trainer. The file is
byte-for-byte preserved; no training logic or paths were changed for this release.
Original SHA-256: `07a53c28262cfcda90ee1a0d37facc82ad6a4de4d5c448305f4ae69202836ca9`.

Validation for this recovery was source-hash equality, Python syntax checking,
and lookup-table shape/range checks. The original GPU experiment was not rerun
as part of publication; this is not a newly validated portable training package.

## Objective and sweep

The script jointly trains the MERT-feature tokenizer head and the acoustic
decoder LoRA. Real-audio flow loss backpropagates through straight-through token
embeddings. A frozen VAE decoder adds log-mel, multi-resolution STFT, and stereo
width losses on an estimated clean latent crop. Minted examples supply head
soft-label CE and 25% of decoder flow windows.

The same script runs every audio-loss variant:

| Run name | `AUX_W` |
| --- | ---: |
| joint_v6 | 0.3 |
| joint_v7 | 1.0 |
| joint_v8 | 2.0 |
| joint_v9 | 4.0 |

**`AUX_W` defaults to zero**, so set it explicitly to enable the waveform loss.
Other defaults: `AUX_TMAX=0.4`, `AUX_FR=150`, `AUX_MARGIN=25`, 512-frame training
windows, and 25 frames/second. The waveform term applies to real windows only
when sampled flow time is at most `AUX_TMAX`. It can also be skipped when the
original audio cannot be resolved or the requested crop is too short.

## Original environment and required inputs

Use the older `yue2-infer` runtime described in the repository README (commit
`92a73cc7`, original Python 3.12 / torch 2.10 / CUDA 12.8 environment), plus numpy,
soundfile, torchaudio, and ffmpeg. Current YuE2 APIs may differ. Do not replace a
working modern training environment with these older dependencies.

The original script has hard-coded paths. Recreate them or edit a working copy:

- `/workspace/hf/hub/models--m-a-p--YuE2-3B/snapshots/...` and the corresponding
  `YuE2-Vae` snapshot. Both must already exist. The script takes the first glob
  match; isolate one intended snapshot to avoid accidental version selection.
- `/workspace/tok/full/feats/<minted_id>.npy`: minted MERT L20 features, either
  `[T,1024]` or the legacy `[4,T,1024]` layout (index 3 is selected).
- `/workspace/yue2-corpus/tracks/<minted_id>/`: `semantic.npy`, `latent.npy`, and
  `request.json` with `style`, `lyrics`, and `seed`. The small AR regularizer pack
  alone is insufficient for this trainer: full features and latents are needed.
- Copy `assets/sem_nbr_idx.npy` and `assets/sem_nbr_cos.npy` from this repository
  to `/workspace/tok/full/`. These are semantic-token neighbor lookup tables,
  not song data. Their exact hashes are in `joint_v6_release.json`.
- `RP` points to prepared real recordings. Each child directory needs `mert.npy`
  `[T,1024]`, `lat.npy` `[T,64]`, and integer `prefix.npy`. Training songs require
  at least 512 frames. Set `HOLD` to existing prepared directory names, separated
  by commas; supply at least one sufficiently long held-out recording.
- Edit `audio_path()` to map each prepared name to its exact original audio.
  The historical mapping recognizes `coheed__<name>` and other `<tag>__<name>`
  layouts. Without a working mapping the auxiliary audio loss may silently
  disappear. Check nonzero `aux` values on eligible real updates; legitimate
  zero values also occur for minted windows and larger flow times.
- Create `/workspace/tok/full/listen_real/` for final renders. Run names should
  be unique: this historical script does not implement safe optimizer resume.

## Checkpoint format and example invocation

Unlike the other adapted scripts in this repository, the recovered original
calls `torch.load` and expects `.pt` dictionaries. For a safetensors release,
convert a trusted checkpoint with the existing `ckpt_io.py` helper first:

```python
import torch
from ckpt_io import load_ckpt
torch.save(load_ckpt("tokenizer_head_v5_30k.safetensors"), "head_init.pt")
```

The historical sweep started from the pretrained head corresponding to
`tokenizer_head_v5_30k.safetensors` and `nar_lora_joint_v4.pt`. All variants used
the same initial weights, rather than chaining v6 into v7 into v8 into v9.
For an already prepared custom dataset, after resolving the paths above:

```sh
mkdir -p /workspace/tok/full/listen_real
RP=/path/to/prepared_real HOLD=heldout_song_1,heldout_song_2 \
MINTED_CAP=4000 AUX_W=4.0 AUX_TMAX=0.4 AUX_FR=150 \
python scripts/joint_v6.py my_audio_loss_run 3000 1 1 \
  /path/to/head_init.pt /path/to/nar_lora_joint_v4.pt 32
```

This example is a historical recipe, not a claim that 3,000 steps or those
hyperparameters are optimal for a new dataset. `best` is selected by held-out
latent flow loss, not the auxiliary spectral metric or a listening score.
The outputs include head/LoRA best and last weights and held-out reconstructions;
they are not new-song artist-conditioning demonstrations.
