|
Download reports/nova-qwen21-latent-training-20260921.md from ntc-ai/model-glue-nova-qwen21-experimental: direct link, hf CLI and curl.
- Browser
- Download file 13.8 kB
-
https://huggingface.co/ntc-ai/model-glue-nova-qwen21-experimental/resolve/main/reports/nova-qwen21-latent-training-20260921.md
- Command line
-
hf download hf://ntc-ai/model-glue-nova-qwen21-experimental/reports/nova-qwen21-latent-training-20260921.md
-
curl -L -o nova-qwen21-latent-training-20260921.md https://huggingface.co/ntc-ai/model-glue-nova-qwen21-experimental/resolve/main/reports/nova-qwen21-latent-training-20260921.md
13.8 kB
| # Silent latent teacher fitting and fast prompt expansion | |
| The user selected the small teacher-fitting experiment before further capacity | |
| changes, retained the original learning rates, and requested faster prompt data. | |
| The latest clarification permits the full decoder→encoder teacher during silent | |
| target preparation. “Low step size” means fewer denoising steps per prompt. | |
| Training and matched training-set audits consume cached latents only. Training | |
| images are neither saved nor reviewed. Separate decoded held-out validation | |
| continues to measure output quality and automatically publish its samples. The | |
| reserved test remains closed. | |
| ## Preserved starting point | |
| The old 180k supervisor was deliberately stopped and saved complete recovery | |
| states: forward step **141,580**, reverse **167,582**. Its best decoded validation | |
| LPIPS remains **0.13630344** at forward step 141,000 and **0.11997008** at reverse | |
| step 120,000. Its fixed controls never ran. The old ledger reports `failed` | |
| because stopping raises after saving; all its workers exited normally from the | |
| supervisor's perspective. Snapshots and inference exports remain intact. | |
| The matched initial EMA audit uses identical latent measurements across splits: | |
| | Direction | All native train NMSE | Native validation NMSE | Validation/train | | |
| |---|---:|---:|---:| | |
| | Nova → Qwen | 0.05835241 | 0.25905055 | 4.44× | | |
| | Qwen → Nova | 0.02427064 | 0.06202652 | 2.56× | | |
| This establishes a latent generalization gap without decoding training examples. | |
| It does not establish a decoded-quality gap of the same magnitude. | |
| ## Objective experiment | |
| Each direction forks the same complete recovery snapshot into two 6,000-update | |
| arms using the same deterministic, kind-balanced 64-row subset. One retains the | |
| paired-error adversarial objective; the other uses unit normalized teacher-latent | |
| MSE plus the existing raw-particle VIC regularizer. The architecture and learning | |
| rates are unchanged. Full model, EMA, critic, optimizer, calibration and RNG state | |
| are restored, and paired-error noise uses absolute parent-plus-local steps. | |
| Both objectives completed. [The compact numeric report](nova-qwen21-latent-probe-20260921.json) | |
| checks identical parent checkpoint hashes, initial model/critic/RNG/data hashes, | |
| subset identities, step-zero measurements and update budgets. All four compact | |
| exports pass bitwise reload/repeat and unchanged-RNG checks. | |
| | Direction | Initial subset NMSE | GAN at 6k | Supervised at 6k | Supervised vs GAN | | |
| |---|---:|---:|---:|---:| | |
| | Nova → Qwen | 0.05333445 | 0.06556350 | 0.03814053 | −41.8% | | |
| | Qwen → Nova | 0.02572833 | 0.02839172 | 0.02183387 | −23.1% | | |
| Native held-out NMSE improves from 0.25905055 to **0.24696433** forward and | |
| 0.06202652 to **0.05769736** reverse under supervision. GAN ends at 0.34561071 | |
| and 0.07527905 respectively. The 6k supervised phases, including audits and | |
| export verification, take 183.45 / 297.82 seconds versus 312.66 / 455.61 seconds | |
| for GAN: observed 1.70× / 1.53× speedups. Initial codec/model loading is outside | |
| these trainer-internal durations. This is not an isolated hardware microbenchmark. | |
| These results support direct latent supervision for the expanded run. The subset | |
| is still not closely interpolated, so capacity limits remain unresolved. Inherited | |
| GAN optimizer moments also affect the short supervised fork. This experiment | |
| does not demonstrate image-quality improvements or convergence to the teacher; | |
| the longer run must establish those with held-out decoded measurements. | |
| The comparison matches starting states, rows and update budget. The objectives | |
| consume randomness differently, so their subsequent minibatch sequences are not | |
| identical. This is a diagnostic of fitting behavior, not a multi-seed causal or | |
| particle-advantage claim. The probe uses latent-only validation and exact exported | |
| forward checks; it does not decode its training or validation rows. | |
| Reproduce configuration selection with: | |
| ```bash | |
| .venv/bin/python -m scripts.prepare_nova_latent_experiment | |
| CUDA_VISIBLE_DEVICES=0,1 \ | |
| PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src \ | |
| .venv/bin/python -u -m scripts.run_bridge_coverage \ | |
| --root artifacts/runs/bridge-nova-qwen21-latent-probe-20260921 \ | |
| --plan configs/bridge-nova-qwen21-latent-probe-queue-20260921.json \ | |
| --no-time-cap | |
| ``` | |
| The active launch additionally carries the stopped prerequisite's compute ledger. | |
| GPU 1 is forward and GPU 0 reverse. Each command requires an unused output root; | |
| preserve prior experiments when reproducing. | |
| ## Throughput changes | |
| Immutable validation teacher RGB and teacher/reference LPIPS are cached on CPU. | |
| On all 176 forward validation rows, scalar evaluation took **86.36 s**, batch-four | |
| cold evaluation **83.05 s**, and warm evaluation **40.39 s**: **2.14× faster** than | |
| the scalar baseline. Cold/warm metrics match exactly for every row and global RNG | |
| is unchanged. Scalar/batch-four mean LPIPS differs by 0.00000285, with maximum | |
| per-row difference 0.001407, within the predeclared BF16 bounds. The cache preserves | |
| native channels-last layout; otherwise teacher-relative LPIPS changes slightly. | |
| Peak batch-four allocated GPU memory was 7.53 GB; CPU cache use is 554 MB. | |
| The corresponding all-row reverse check takes 32.91 s scalar, 35.42 s batch-four | |
| cold and **17.48 s cached** (1.88× scalar speed). All cold/warm row metrics are | |
| exactly equal and RNG is unchanged. Reverse scalar/batched LPIPS differs by | |
| 0.00015174 mean / 0.00288764 maximum per row, within the same declared bounds. | |
| The old reverse best remeasures at **0.12012182** under batch four rather than | |
| 0.11997008 scalar; comparisons must account for this BF16 batching shift. Batching | |
| alone does not deliver the improvement. The long run uses batch four consistently, | |
| and its movable/fixed controls use the identical evaluation path. | |
| Training caches exact native source-prefix features and prepared noise features. | |
| Their use is checked against ordinary forward predictions. Inference continues | |
| to run the entire bridge forward; none of these caches is exported or retrieved. | |
| Data expansion reuses the original checked train/validation shards, caches prompt | |
| embeddings and keeps native denoisers resident. It computes offline teacher | |
| targets after capturing native source latents. The first batch-four teacher | |
| check failed the predeclared numerical bounds: sampler-coordinate relative NMSE | |
| was 0.00321991 forward and 0.00011402 reverse, with maximum differences 1.11184 | |
| and 0.07687. No case was committed. The run therefore uses **single-example | |
| teacher passes**, preserving the original target semantics rather than relaxing | |
| the bounds. Batched target preparation remains available only for codecs that | |
| pass the scalar comparison. Decoded validation batching passed its separate checks. | |
| New trajectories start at **8 | |
| denoising steps**, capturing steps 1, 3, 5, 7 and the finished state. The original | |
| 40-step rows retain their metadata and remain in the pool. Numerical scalar/batch | |
| teacher checks are implementation checks; no output-dependent example selection | |
| occurs. Both directions are committed atomically at case boundaries. | |
| The deterministic prompt selection adds 128 unused training prompts from the | |
| existing curated pool, with independent seeds. This doubles distinct training | |
| trajectories from 128 to 256 and increases cached rows from 896 to 1,536 per | |
| direction. The 16 held-out prompts and their 176 validation rows are inherited | |
| unchanged. Uniform row sampling gives each old trajectory seven entries and each | |
| new trajectory five, rather than weighting trajectories equally. Future 16/32-step | |
| expansions can use the same generic expansion tool; | |
| they should be triggered by held-out numerical evidence, without selecting or | |
| reviewing training images. | |
| ## Verification and controls | |
| The final clean-checkout verification suite passed **92 tests**. Actual GPU reverse | |
| fresh movable/fixed calibration smoke checks passed: initial model/critic/RNG/data | |
| hashes match, native weights are unchanged, cached/direct prefix predictions are | |
| bitwise equal, and exported inference is bitwise repeatable without consuming RNG. | |
| After two updates, movable cloud displacement was 0.08012274 and fixed displacement | |
| zero. No image files were generated by the smoke run. | |
| This caught and fixed a real initialization bug: fixed particles are registered | |
| as a buffer, so copying every buffer from a trained calibration checkpoint would | |
| accidentally import its trained cloud only into the fixed arm. Calibration import | |
| now names immutable unit buffers explicitly and verifies learned weights and | |
| particles stayed unchanged. Fresh long-run controls therefore share initial | |
| learned weights, particles and calibration. | |
| Detailed development evidence is under | |
| `artifacts/verification/nova-qwen21-latent-speed-20260921/`. Runtime datasets, | |
| checkpoints, prompt embeddings and verification images are ignored artifacts. | |
| ## Bootstrapping another pair without a trained parent | |
| `scripts.prepare_image_bridge_calibration` prepares immutable latent/feature | |
| units offline from pinned training latents. Full decode/encode computation occurs | |
| in this preparation stage only, without saving or reviewing training images. | |
| The compact artifact contains calibration buffers, their checksum, the native | |
| weight checksum and dataset/codec provenance. It contains no learned weights or | |
| optimizer state. The trainer accepts it through `--calibration-from` and checks | |
| native weights before importing the buffers. | |
| For example, after preparing the pinned dataset: | |
| ```bash | |
| CUDA_VISIBLE_DEVICES=0 \ | |
| PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src \ | |
| .venv/bin/python -m scripts.prepare_image_bridge_calibration \ | |
| --config configs/bridge-nova-qwen21-latent300-20260921.json \ | |
| --arm reverse --output artifacts/calibration/nova-qwen21-reverse.pt | |
| ``` | |
| Run the forward counterpart on GPU 1 with `--arm forward`. Calibration refuses | |
| held-out rows before native processing; it loads only training shards. The actual | |
| GPU smoke used 896 training latents for affine/prefix units and four rows for the | |
| head check (the production default is 64), yielding a 19,661-byte artifact. Fresh | |
| movable/fixed imports had identical initial hashes, exact cached-prefix and | |
| exported-forward checks, unchanged native weights, and zero saved images. The | |
| fixed cloud stayed unchanged; the movable cloud moved. The active longer campaign | |
| retains its previously verified recovery-buffer calibration, so this additional | |
| bootstrap does not change that experiment's initialization. | |
| ## Completed data preparation and active longer run | |
| All 128 new prompts completed. Both directions have **1,536 training rows and | |
| 176 validation rows**. Every new shard checksum/schema was checked; training | |
| shards contain source/target latents and noise fraction only. Inherited validation | |
| rows and checksums match exactly. Scalar teacher repeat checks are bitwise equal. | |
| See [the data evidence](nova-qwen21-fast-data-evidence-20260921.json). | |
| After initial loading, new cases average **6.085 s** (median 5.178 s), compared | |
| with the historical completed-case interval mean of **27.234 s**: **4.48× faster | |
| per prompt** across the full batch. Both statistics exclude the first case and | |
| initial loading. The old timing is reconstructed from immutable shard creation | |
| times; this is an observational comparison. Earlier partial-batch estimates were | |
| faster than the final mean. Prompt contexts are cached for resumptions. | |
| The active supervisor is **2194755**, root | |
| `artifacts/runs/bridge-nova-qwen21-latent300-20260921`. Data preparation exited | |
| successfully; forward trainer **2215023** runs on GPU 1 and reverse **2215024** | |
| on GPU 0. Both have advanced beyond startup. They use all 1,536 training rows, | |
| unchanged learning rates and fresh learned weights/optimizers. Both native feature | |
| caches are enabled and verified bitwise against ordinary predictions. Source | |
| for the trainers is **f1af678**; the pipeline launch is recorded under **3c8eae1** | |
| before the additive offline-bootstrap support was committed. | |
| Each direction targets **300,000 updates**, with decoded validation and matched | |
| latent audits every 6,000. Best checkpoints are selected by validation LPIPS. | |
| Fresh matched fixed controls run next, followed by finished/round-trip/switched | |
| qualification and the package report. Old best checkpoints remain preserved. | |
| The [first trained decoded validation](nova-qwen21-latent300-first-validation-20260921.json) | |
| at **6,000 updates** improves both directions relative to the preserved best, | |
| remeasured with the same batch-four evaluator: | |
| | Direction | Previous best LPIPS, batch four | New 6k LPIPS | Relative change | | |
| |---|---:|---:|---:| | |
| | Nova → Qwen | 0.13630629 | **0.11840734** | −13.13% | | |
| | Qwen → Nova | 0.12012182 | **0.11876335** | −1.13% | | |
| Both new checkpoints are selected at 6k. All 176 teacher-cache entries are reused; | |
| the actual validations take 29.81 / 21.66 seconds while the opposite direction | |
| is training. The new training recipe changes objective, coverage and initialization | |
| together; these figures do not isolate their individual contributions. This is | |
| one early validation milestone, not a convergence, particle-advantage or | |
| distribution-readiness claim. The runs continue, and the test remains unopened. | |
| The dashboard was restarted as **2179728** on | |
| [port 8779](http://100.90.104.57:8779/#nova-curves). The supervised campaign is | |
| selected by default and keeps earlier campaigns available. Its first decoded | |
| validation samples publish automatically; no training images are published. | |
| Live browser checks passed on desktop and 390px mobile: both running directions, | |
| supervised and legacy GAN metrics, queued fixed controls, refresh persistence, | |
| and all 12 SD/Sana histories. No JavaScript errors or horizontal overflow occurred. | |