model-glue-nova-qwen21-experimental / reports /nova-qwen21-latent-training-20260921.md
ntc-ai's picture
Archive Nova Qwen bridge checkpoints and measured combined sampling
1f167a1 verified
|
Raw
History Blame Contribute Delete
13.8 kB

Silent latent teacher fitting and fast prompt expansion

The user selected the small teacher-fitting experiment before further capacity changes, retained the original learning rates, and requested faster prompt data. The latest clarification permits the full decoder→encoder teacher during silent target preparation. “Low step size” means fewer denoising steps per prompt.

Training and matched training-set audits consume cached latents only. Training images are neither saved nor reviewed. Separate decoded held-out validation continues to measure output quality and automatically publish its samples. The reserved test remains closed.

Preserved starting point

The old 180k supervisor was deliberately stopped and saved complete recovery states: forward step 141,580, reverse 167,582. Its best decoded validation LPIPS remains 0.13630344 at forward step 141,000 and 0.11997008 at reverse step 120,000. Its fixed controls never ran. The old ledger reports failed because stopping raises after saving; all its workers exited normally from the supervisor's perspective. Snapshots and inference exports remain intact.

The matched initial EMA audit uses identical latent measurements across splits:

Direction All native train NMSE Native validation NMSE Validation/train
Nova → Qwen 0.05835241 0.25905055 4.44×
Qwen → Nova 0.02427064 0.06202652 2.56×

This establishes a latent generalization gap without decoding training examples. It does not establish a decoded-quality gap of the same magnitude.

Objective experiment

Each direction forks the same complete recovery snapshot into two 6,000-update arms using the same deterministic, kind-balanced 64-row subset. One retains the paired-error adversarial objective; the other uses unit normalized teacher-latent MSE plus the existing raw-particle VIC regularizer. The architecture and learning rates are unchanged. Full model, EMA, critic, optimizer, calibration and RNG state are restored, and paired-error noise uses absolute parent-plus-local steps.

Both objectives completed. The compact numeric report checks identical parent checkpoint hashes, initial model/critic/RNG/data hashes, subset identities, step-zero measurements and update budgets. All four compact exports pass bitwise reload/repeat and unchanged-RNG checks.

Direction Initial subset NMSE GAN at 6k Supervised at 6k Supervised vs GAN
Nova → Qwen 0.05333445 0.06556350 0.03814053 −41.8%
Qwen → Nova 0.02572833 0.02839172 0.02183387 −23.1%

Native held-out NMSE improves from 0.25905055 to 0.24696433 forward and 0.06202652 to 0.05769736 reverse under supervision. GAN ends at 0.34561071 and 0.07527905 respectively. The 6k supervised phases, including audits and export verification, take 183.45 / 297.82 seconds versus 312.66 / 455.61 seconds for GAN: observed 1.70× / 1.53× speedups. Initial codec/model loading is outside these trainer-internal durations. This is not an isolated hardware microbenchmark.

These results support direct latent supervision for the expanded run. The subset is still not closely interpolated, so capacity limits remain unresolved. Inherited GAN optimizer moments also affect the short supervised fork. This experiment does not demonstrate image-quality improvements or convergence to the teacher; the longer run must establish those with held-out decoded measurements.

The comparison matches starting states, rows and update budget. The objectives consume randomness differently, so their subsequent minibatch sequences are not identical. This is a diagnostic of fitting behavior, not a multi-seed causal or particle-advantage claim. The probe uses latent-only validation and exact exported forward checks; it does not decode its training or validation rows.

Reproduce configuration selection with:

.venv/bin/python -m scripts.prepare_nova_latent_experiment
CUDA_VISIBLE_DEVICES=0,1 \
PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src \
.venv/bin/python -u -m scripts.run_bridge_coverage \
  --root artifacts/runs/bridge-nova-qwen21-latent-probe-20260921 \
  --plan configs/bridge-nova-qwen21-latent-probe-queue-20260921.json \
  --no-time-cap

The active launch additionally carries the stopped prerequisite's compute ledger. GPU 1 is forward and GPU 0 reverse. Each command requires an unused output root; preserve prior experiments when reproducing.

Throughput changes

Immutable validation teacher RGB and teacher/reference LPIPS are cached on CPU. On all 176 forward validation rows, scalar evaluation took 86.36 s, batch-four cold evaluation 83.05 s, and warm evaluation 40.39 s: 2.14× faster than the scalar baseline. Cold/warm metrics match exactly for every row and global RNG is unchanged. Scalar/batch-four mean LPIPS differs by 0.00000285, with maximum per-row difference 0.001407, within the predeclared BF16 bounds. The cache preserves native channels-last layout; otherwise teacher-relative LPIPS changes slightly. Peak batch-four allocated GPU memory was 7.53 GB; CPU cache use is 554 MB.

The corresponding all-row reverse check takes 32.91 s scalar, 35.42 s batch-four cold and 17.48 s cached (1.88× scalar speed). All cold/warm row metrics are exactly equal and RNG is unchanged. Reverse scalar/batched LPIPS differs by 0.00015174 mean / 0.00288764 maximum per row, within the same declared bounds. The old reverse best remeasures at 0.12012182 under batch four rather than 0.11997008 scalar; comparisons must account for this BF16 batching shift. Batching alone does not deliver the improvement. The long run uses batch four consistently, and its movable/fixed controls use the identical evaluation path.

Training caches exact native source-prefix features and prepared noise features. Their use is checked against ordinary forward predictions. Inference continues to run the entire bridge forward; none of these caches is exported or retrieved.

Data expansion reuses the original checked train/validation shards, caches prompt embeddings and keeps native denoisers resident. It computes offline teacher targets after capturing native source latents. The first batch-four teacher check failed the predeclared numerical bounds: sampler-coordinate relative NMSE was 0.00321991 forward and 0.00011402 reverse, with maximum differences 1.11184 and 0.07687. No case was committed. The run therefore uses single-example teacher passes, preserving the original target semantics rather than relaxing the bounds. Batched target preparation remains available only for codecs that pass the scalar comparison. Decoded validation batching passed its separate checks. New trajectories start at 8 denoising steps, capturing steps 1, 3, 5, 7 and the finished state. The original 40-step rows retain their metadata and remain in the pool. Numerical scalar/batch teacher checks are implementation checks; no output-dependent example selection occurs. Both directions are committed atomically at case boundaries.

The deterministic prompt selection adds 128 unused training prompts from the existing curated pool, with independent seeds. This doubles distinct training trajectories from 128 to 256 and increases cached rows from 896 to 1,536 per direction. The 16 held-out prompts and their 176 validation rows are inherited unchanged. Uniform row sampling gives each old trajectory seven entries and each new trajectory five, rather than weighting trajectories equally. Future 16/32-step expansions can use the same generic expansion tool; they should be triggered by held-out numerical evidence, without selecting or reviewing training images.

Verification and controls

The final clean-checkout verification suite passed 92 tests. Actual GPU reverse fresh movable/fixed calibration smoke checks passed: initial model/critic/RNG/data hashes match, native weights are unchanged, cached/direct prefix predictions are bitwise equal, and exported inference is bitwise repeatable without consuming RNG. After two updates, movable cloud displacement was 0.08012274 and fixed displacement zero. No image files were generated by the smoke run.

This caught and fixed a real initialization bug: fixed particles are registered as a buffer, so copying every buffer from a trained calibration checkpoint would accidentally import its trained cloud only into the fixed arm. Calibration import now names immutable unit buffers explicitly and verifies learned weights and particles stayed unchanged. Fresh long-run controls therefore share initial learned weights, particles and calibration.

Detailed development evidence is under artifacts/verification/nova-qwen21-latent-speed-20260921/. Runtime datasets, checkpoints, prompt embeddings and verification images are ignored artifacts.

Bootstrapping another pair without a trained parent

scripts.prepare_image_bridge_calibration prepares immutable latent/feature units offline from pinned training latents. Full decode/encode computation occurs in this preparation stage only, without saving or reviewing training images. The compact artifact contains calibration buffers, their checksum, the native weight checksum and dataset/codec provenance. It contains no learned weights or optimizer state. The trainer accepts it through --calibration-from and checks native weights before importing the buffers.

For example, after preparing the pinned dataset:

CUDA_VISIBLE_DEVICES=0 \
PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src \
.venv/bin/python -m scripts.prepare_image_bridge_calibration \
  --config configs/bridge-nova-qwen21-latent300-20260921.json \
  --arm reverse --output artifacts/calibration/nova-qwen21-reverse.pt

Run the forward counterpart on GPU 1 with --arm forward. Calibration refuses held-out rows before native processing; it loads only training shards. The actual GPU smoke used 896 training latents for affine/prefix units and four rows for the head check (the production default is 64), yielding a 19,661-byte artifact. Fresh movable/fixed imports had identical initial hashes, exact cached-prefix and exported-forward checks, unchanged native weights, and zero saved images. The fixed cloud stayed unchanged; the movable cloud moved. The active longer campaign retains its previously verified recovery-buffer calibration, so this additional bootstrap does not change that experiment's initialization.

Completed data preparation and active longer run

All 128 new prompts completed. Both directions have 1,536 training rows and 176 validation rows. Every new shard checksum/schema was checked; training shards contain source/target latents and noise fraction only. Inherited validation rows and checksums match exactly. Scalar teacher repeat checks are bitwise equal. See the data evidence.

After initial loading, new cases average 6.085 s (median 5.178 s), compared with the historical completed-case interval mean of 27.234 s: 4.48× faster per prompt across the full batch. Both statistics exclude the first case and initial loading. The old timing is reconstructed from immutable shard creation times; this is an observational comparison. Earlier partial-batch estimates were faster than the final mean. Prompt contexts are cached for resumptions.

The active supervisor is 2194755, root artifacts/runs/bridge-nova-qwen21-latent300-20260921. Data preparation exited successfully; forward trainer 2215023 runs on GPU 1 and reverse 2215024 on GPU 0. Both have advanced beyond startup. They use all 1,536 training rows, unchanged learning rates and fresh learned weights/optimizers. Both native feature caches are enabled and verified bitwise against ordinary predictions. Source for the trainers is f1af678; the pipeline launch is recorded under 3c8eae1 before the additive offline-bootstrap support was committed.

Each direction targets 300,000 updates, with decoded validation and matched latent audits every 6,000. Best checkpoints are selected by validation LPIPS. Fresh matched fixed controls run next, followed by finished/round-trip/switched qualification and the package report. Old best checkpoints remain preserved.

The first trained decoded validation at 6,000 updates improves both directions relative to the preserved best, remeasured with the same batch-four evaluator:

Direction Previous best LPIPS, batch four New 6k LPIPS Relative change
Nova → Qwen 0.13630629 0.11840734 −13.13%
Qwen → Nova 0.12012182 0.11876335 −1.13%

Both new checkpoints are selected at 6k. All 176 teacher-cache entries are reused; the actual validations take 29.81 / 21.66 seconds while the opposite direction is training. The new training recipe changes objective, coverage and initialization together; these figures do not isolate their individual contributions. This is one early validation milestone, not a convergence, particle-advantage or distribution-readiness claim. The runs continue, and the test remains unopened.

The dashboard was restarted as 2179728 on port 8779. The supervised campaign is selected by default and keeps earlier campaigns available. Its first decoded validation samples publish automatically; no training images are published. Live browser checks passed on desktop and 390px mobile: both running directions, supervised and legacy GAN metrics, queued fixed controls, refresh persistence, and all 12 SD/Sana histories. No JavaScript errors or horizontal overflow occurred.