model-glue-nova-qwen21-experimental / reports /nova-qwen21-latent-training-20260921.md
ntc-ai's picture
Archive Nova Qwen bridge checkpoints and measured combined sampling
1f167a1 verified
|
Raw History Blame Contribute Delete
13.8 kB
# Silent latent teacher fitting and fast prompt expansion
The user selected the small teacher-fitting experiment before further capacity
changes, retained the original learning rates, and requested faster prompt data.
The latest clarification permits the full decoder→encoder teacher during silent
target preparation. “Low step size” means fewer denoising steps per prompt.
Training and matched training-set audits consume cached latents only. Training
images are neither saved nor reviewed. Separate decoded held-out validation
continues to measure output quality and automatically publish its samples. The
reserved test remains closed.
## Preserved starting point
The old 180k supervisor was deliberately stopped and saved complete recovery
states: forward step **141,580**, reverse **167,582**. Its best decoded validation
LPIPS remains **0.13630344** at forward step 141,000 and **0.11997008** at reverse
step 120,000. Its fixed controls never ran. The old ledger reports `failed`
because stopping raises after saving; all its workers exited normally from the
supervisor's perspective. Snapshots and inference exports remain intact.
The matched initial EMA audit uses identical latent measurements across splits:
| Direction | All native train NMSE | Native validation NMSE | Validation/train |
|---|---:|---:|---:|
| Nova → Qwen | 0.05835241 | 0.25905055 | 4.44× |
| Qwen → Nova | 0.02427064 | 0.06202652 | 2.56× |
This establishes a latent generalization gap without decoding training examples.
It does not establish a decoded-quality gap of the same magnitude.
## Objective experiment
Each direction forks the same complete recovery snapshot into two 6,000-update
arms using the same deterministic, kind-balanced 64-row subset. One retains the
paired-error adversarial objective; the other uses unit normalized teacher-latent
MSE plus the existing raw-particle VIC regularizer. The architecture and learning
rates are unchanged. Full model, EMA, critic, optimizer, calibration and RNG state
are restored, and paired-error noise uses absolute parent-plus-local steps.
Both objectives completed. [The compact numeric report](nova-qwen21-latent-probe-20260921.json)
checks identical parent checkpoint hashes, initial model/critic/RNG/data hashes,
subset identities, step-zero measurements and update budgets. All four compact
exports pass bitwise reload/repeat and unchanged-RNG checks.
| Direction | Initial subset NMSE | GAN at 6k | Supervised at 6k | Supervised vs GAN |
|---|---:|---:|---:|---:|
| Nova → Qwen | 0.05333445 | 0.06556350 | 0.03814053 | −41.8% |
| Qwen → Nova | 0.02572833 | 0.02839172 | 0.02183387 | −23.1% |
Native held-out NMSE improves from 0.25905055 to **0.24696433** forward and
0.06202652 to **0.05769736** reverse under supervision. GAN ends at 0.34561071
and 0.07527905 respectively. The 6k supervised phases, including audits and
export verification, take 183.45 / 297.82 seconds versus 312.66 / 455.61 seconds
for GAN: observed 1.70× / 1.53× speedups. Initial codec/model loading is outside
these trainer-internal durations. This is not an isolated hardware microbenchmark.
These results support direct latent supervision for the expanded run. The subset
is still not closely interpolated, so capacity limits remain unresolved. Inherited
GAN optimizer moments also affect the short supervised fork. This experiment
does not demonstrate image-quality improvements or convergence to the teacher;
the longer run must establish those with held-out decoded measurements.
The comparison matches starting states, rows and update budget. The objectives
consume randomness differently, so their subsequent minibatch sequences are not
identical. This is a diagnostic of fitting behavior, not a multi-seed causal or
particle-advantage claim. The probe uses latent-only validation and exact exported
forward checks; it does not decode its training or validation rows.
Reproduce configuration selection with:
```bash
.venv/bin/python -m scripts.prepare_nova_latent_experiment
CUDA_VISIBLE_DEVICES=0,1 \
PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src \
.venv/bin/python -u -m scripts.run_bridge_coverage \
--root artifacts/runs/bridge-nova-qwen21-latent-probe-20260921 \
--plan configs/bridge-nova-qwen21-latent-probe-queue-20260921.json \
--no-time-cap
```
The active launch additionally carries the stopped prerequisite's compute ledger.
GPU 1 is forward and GPU 0 reverse. Each command requires an unused output root;
preserve prior experiments when reproducing.
## Throughput changes
Immutable validation teacher RGB and teacher/reference LPIPS are cached on CPU.
On all 176 forward validation rows, scalar evaluation took **86.36 s**, batch-four
cold evaluation **83.05 s**, and warm evaluation **40.39 s**: **2.14× faster** than
the scalar baseline. Cold/warm metrics match exactly for every row and global RNG
is unchanged. Scalar/batch-four mean LPIPS differs by 0.00000285, with maximum
per-row difference 0.001407, within the predeclared BF16 bounds. The cache preserves
native channels-last layout; otherwise teacher-relative LPIPS changes slightly.
Peak batch-four allocated GPU memory was 7.53 GB; CPU cache use is 554 MB.
The corresponding all-row reverse check takes 32.91 s scalar, 35.42 s batch-four
cold and **17.48 s cached** (1.88× scalar speed). All cold/warm row metrics are
exactly equal and RNG is unchanged. Reverse scalar/batched LPIPS differs by
0.00015174 mean / 0.00288764 maximum per row, within the same declared bounds.
The old reverse best remeasures at **0.12012182** under batch four rather than
0.11997008 scalar; comparisons must account for this BF16 batching shift. Batching
alone does not deliver the improvement. The long run uses batch four consistently,
and its movable/fixed controls use the identical evaluation path.
Training caches exact native source-prefix features and prepared noise features.
Their use is checked against ordinary forward predictions. Inference continues
to run the entire bridge forward; none of these caches is exported or retrieved.
Data expansion reuses the original checked train/validation shards, caches prompt
embeddings and keeps native denoisers resident. It computes offline teacher
targets after capturing native source latents. The first batch-four teacher
check failed the predeclared numerical bounds: sampler-coordinate relative NMSE
was 0.00321991 forward and 0.00011402 reverse, with maximum differences 1.11184
and 0.07687. No case was committed. The run therefore uses **single-example
teacher passes**, preserving the original target semantics rather than relaxing
the bounds. Batched target preparation remains available only for codecs that
pass the scalar comparison. Decoded validation batching passed its separate checks.
New trajectories start at **8
denoising steps**, capturing steps 1, 3, 5, 7 and the finished state. The original
40-step rows retain their metadata and remain in the pool. Numerical scalar/batch
teacher checks are implementation checks; no output-dependent example selection
occurs. Both directions are committed atomically at case boundaries.
The deterministic prompt selection adds 128 unused training prompts from the
existing curated pool, with independent seeds. This doubles distinct training
trajectories from 128 to 256 and increases cached rows from 896 to 1,536 per
direction. The 16 held-out prompts and their 176 validation rows are inherited
unchanged. Uniform row sampling gives each old trajectory seven entries and each
new trajectory five, rather than weighting trajectories equally. Future 16/32-step
expansions can use the same generic expansion tool;
they should be triggered by held-out numerical evidence, without selecting or
reviewing training images.
## Verification and controls
The final clean-checkout verification suite passed **92 tests**. Actual GPU reverse
fresh movable/fixed calibration smoke checks passed: initial model/critic/RNG/data
hashes match, native weights are unchanged, cached/direct prefix predictions are
bitwise equal, and exported inference is bitwise repeatable without consuming RNG.
After two updates, movable cloud displacement was 0.08012274 and fixed displacement
zero. No image files were generated by the smoke run.
This caught and fixed a real initialization bug: fixed particles are registered
as a buffer, so copying every buffer from a trained calibration checkpoint would
accidentally import its trained cloud only into the fixed arm. Calibration import
now names immutable unit buffers explicitly and verifies learned weights and
particles stayed unchanged. Fresh long-run controls therefore share initial
learned weights, particles and calibration.
Detailed development evidence is under
`artifacts/verification/nova-qwen21-latent-speed-20260921/`. Runtime datasets,
checkpoints, prompt embeddings and verification images are ignored artifacts.
## Bootstrapping another pair without a trained parent
`scripts.prepare_image_bridge_calibration` prepares immutable latent/feature
units offline from pinned training latents. Full decode/encode computation occurs
in this preparation stage only, without saving or reviewing training images.
The compact artifact contains calibration buffers, their checksum, the native
weight checksum and dataset/codec provenance. It contains no learned weights or
optimizer state. The trainer accepts it through `--calibration-from` and checks
native weights before importing the buffers.
For example, after preparing the pinned dataset:
```bash
CUDA_VISIBLE_DEVICES=0 \
PYTHONPATH=.venv/qwen21-deps:artifacts/vendor/diffusers-qwen21/src \
.venv/bin/python -m scripts.prepare_image_bridge_calibration \
--config configs/bridge-nova-qwen21-latent300-20260921.json \
--arm reverse --output artifacts/calibration/nova-qwen21-reverse.pt
```
Run the forward counterpart on GPU 1 with `--arm forward`. Calibration refuses
held-out rows before native processing; it loads only training shards. The actual
GPU smoke used 896 training latents for affine/prefix units and four rows for the
head check (the production default is 64), yielding a 19,661-byte artifact. Fresh
movable/fixed imports had identical initial hashes, exact cached-prefix and
exported-forward checks, unchanged native weights, and zero saved images. The
fixed cloud stayed unchanged; the movable cloud moved. The active longer campaign
retains its previously verified recovery-buffer calibration, so this additional
bootstrap does not change that experiment's initialization.
## Completed data preparation and active longer run
All 128 new prompts completed. Both directions have **1,536 training rows and
176 validation rows**. Every new shard checksum/schema was checked; training
shards contain source/target latents and noise fraction only. Inherited validation
rows and checksums match exactly. Scalar teacher repeat checks are bitwise equal.
See [the data evidence](nova-qwen21-fast-data-evidence-20260921.json).
After initial loading, new cases average **6.085 s** (median 5.178 s), compared
with the historical completed-case interval mean of **27.234 s**: **4.48× faster
per prompt** across the full batch. Both statistics exclude the first case and
initial loading. The old timing is reconstructed from immutable shard creation
times; this is an observational comparison. Earlier partial-batch estimates were
faster than the final mean. Prompt contexts are cached for resumptions.
The active supervisor is **2194755**, root
`artifacts/runs/bridge-nova-qwen21-latent300-20260921`. Data preparation exited
successfully; forward trainer **2215023** runs on GPU 1 and reverse **2215024**
on GPU 0. Both have advanced beyond startup. They use all 1,536 training rows,
unchanged learning rates and fresh learned weights/optimizers. Both native feature
caches are enabled and verified bitwise against ordinary predictions. Source
for the trainers is **f1af678**; the pipeline launch is recorded under **3c8eae1**
before the additive offline-bootstrap support was committed.
Each direction targets **300,000 updates**, with decoded validation and matched
latent audits every 6,000. Best checkpoints are selected by validation LPIPS.
Fresh matched fixed controls run next, followed by finished/round-trip/switched
qualification and the package report. Old best checkpoints remain preserved.
The [first trained decoded validation](nova-qwen21-latent300-first-validation-20260921.json)
at **6,000 updates** improves both directions relative to the preserved best,
remeasured with the same batch-four evaluator:
| Direction | Previous best LPIPS, batch four | New 6k LPIPS | Relative change |
|---|---:|---:|---:|
| Nova → Qwen | 0.13630629 | **0.11840734** | −13.13% |
| Qwen → Nova | 0.12012182 | **0.11876335** | −1.13% |
Both new checkpoints are selected at 6k. All 176 teacher-cache entries are reused;
the actual validations take 29.81 / 21.66 seconds while the opposite direction
is training. The new training recipe changes objective, coverage and initialization
together; these figures do not isolate their individual contributions. This is
one early validation milestone, not a convergence, particle-advantage or
distribution-readiness claim. The runs continue, and the test remains unopened.
The dashboard was restarted as **2179728** on
[port 8779](http://100.90.104.57:8779/#nova-curves). The supervised campaign is
selected by default and keeps earlier campaigns available. Its first decoded
validation samples publish automatically; no training images are published.
Live browser checks passed on desktop and 390px mobile: both running directions,
supervised and legacy GAN metrics, queued fixed controls, refresh persistence,
and all 12 SD/Sana histories. No JavaScript errors or horizontal overflow occurred.