model-glue-nova-qwen21-experimental / reports /nova-qwen21-extension-20260921.md
ntc-ai's picture
Archive Nova Qwen bridge checkpoints and measured combined sampling
1f167a1 verified
|
Raw History Blame Contribute Delete
2.59 kB

Nova/Qwen native-head continuation to 180k

The user requested longer training while validation continued to improve. Nova → Qwen LPIPS fell from 0.248844 at 30k to 0.204068 at 36k. Qwen → Nova fell from 0.145471 at 21k to 0.131442 at 30k. These are validation measurements, not a claim that the bridges are ready for distribution.

Both directions now target 180,000 updates. The configuration is identical to the 36k native-head recipe except for training.steps. The queue then trains fresh matched fixed-cloud controls to 180k and runs the existing selected-pair finished, round-trip and switching qualification. Checkpoint selection remains best validation LPIPS; test stays reserved.

The old supervisor was gracefully stopped. Forward continues from its completed 36,000-step recovery; reverse saved step 30,094, including the 30k validation. Both retain model, EMA, critic, optimizers, particles, normalization and CPU, CUDA and training-generator RNG state. Old checkpoints remain preserved in the head36 root. The stopped reverse exit is intentionally marked failed by the old supervisor; its successor accounts for that ledger's 4,313.640639 GPU-seconds.

Configuration: configs/bridge-nova-qwen21-head180-20260921.json. Plan: configs/bridge-nova-qwen21-head180-queue-20260921.json. Run root: artifacts/runs/bridge-nova-qwen21-head180-20260921. Training source: 6f1e8b2. Launch metadata and the exact preserved checkpoint steps are recorded under the run root. GPU 1 remains forward; GPU 0 reverse.

The explicit stopped-parent option rejects active supervisors or unfinished worker records. A continuation still rejects changes to the recipe, pair, seed, cloud, direction or dataset. Fourteen focused CPU tests passed for native-head training, continuation constraints and compute accounting. Both real recovery states passed the same continuation checks before launch.

Both actual GPU preflights passed native prefix/head equality, optimizer wiring and cached-feature prediction equality. The live browser showed forward at 36,900 and reverse at 30,900, both running toward 180k with inherited validation history intact. All four graphs rendered without JavaScript errors. The run's launch-verification.json records these observations.

Live graphs retain the complete prior history and now show a 180k target. Validation and completed image samples continue to publish every 3,000 updates. The default campaign remains the current run, with the initial horizon and original direct-map campaign available for comparison.