Download reports/nova-qwen21-convergence-strategy-20260921.md from ntc-ai/model-glue-nova-qwen21-experimental: direct link, hf CLI and curl.
- Browser
- Download file 18.7 kB
-
https://huggingface.co/ntc-ai/model-glue-nova-qwen21-experimental/resolve/main/reports/nova-qwen21-convergence-strategy-20260921.md
- Command line
-
hf download hf://ntc-ai/model-glue-nova-qwen21-experimental/reports/nova-qwen21-convergence-strategy-20260921.md
-
curl -L -o nova-qwen21-convergence-strategy-20260921.md https://huggingface.co/ntc-ai/model-glue-nova-qwen21-experimental/resolve/main/reports/nova-qwen21-convergence-strategy-20260921.md
Nova/Qwen convergence strategy, September 21
Three independent research reviews covered capacity/pretrained reuse, optimization, and interface/data/evaluation limits. The synthesis below combines their code audits, primary literature, live metrics and existing validation images. This is a proposed experiment sequence; it does not modify the running 180k campaign or its matched controls.
The best-supported hypothesis is a combination of limited training coverage and a restrictive learned approximation, with constant-rate optimization adding noise. There is no demonstrated codec incompatibility preventing a much more accurate converter. No experiment or cited theorem guarantees that the present small adversarially trained network will reach the full-codec reference.
Evidence and what it establishes
The accompanying measurement snapshot records the actual run paths, read time, quality strata and training diagnostics. At that snapshot:
| Direction | Selected step | Source LPIPS | Student-to-teacher LPIPS | Full-codec source LPIPS |
|---|---|---|---|---|
| Nova → Qwen | 126,000 | 0.138064 | 0.137344 | 0.010601 |
| Qwen → Nova | 120,000 | 0.119970 | 0.098619 | 0.054272 |
Forward's median validation LPIPS fell from 0.18963 over 36k–60k to 0.14715 over 93k–120k. Reverse's corresponding medians were 0.12425 and 0.12037; over 123k–150k its median was 0.12026. Forward still improves, while reverse is nearly flat. A finished forward kingfisher at 126k preserves composition but has clear block and edge distortions absent from the source and full-codec target.
The exact teacher is
T(z) = E_target(quantize(channel_convert(clamp(D_source(z))))), using posterior
mode. It needs only the raw latent already supplied to the learned bridge.
The native decoder remainder consumes the retained source-prefix features, so
those features are sufficient to compute the teacher in principle. Teacher
targets also pass through the recipient head that the learned bridge uses.
The restriction is the compact map inserted between the frozen endpoints.
Native prefix/head, stored-pair, cache and exported-forward checks have passed;
see the original diagnosis.
The full-codec score is a measured reference, not a mathematical lower bound. Student-to-teacher fidelity and source preservation must both be measured. General exact reversibility is impossible across the declared RGBA→RGB compositing operation: alpha is discarded. Quantization and codec compression also limit round trips. Equal-dimensional Gaussian noise transport does not align the two diffusion vector fields, so conversion accuracy and switched generation require separate qualification.
Capacity and pretrained processing
| Property | Forward | Reverse |
|---|---|---|
| Raw input | 16×64×64 | 64×32×32 |
| Number of latent scalars | 65,536 | 65,536 |
| Native feature channels before learned projection | 384 | 1,152 |
| Projected channels | 96 | 96 |
| Learned residual blocks | 4 | 4 |
| Predicted recipient feature channels | 768 | 384 |
| Trainable parameters | 2,217,828 | 2,560,356 |
Pixel shuffle/unshuffle is lossless. Detailed spatial information instead passes through a pointwise 96-channel projection, without an independent raw-latent spatial bypass. This limits the function class but does not prove information loss on the actual data manifold. The source prefix already contains native attention; the whole bridge does not lack global context. The learned global particle condition is only four-dimensional.
The learned map skips source decoder upsampling blocks and recipient encoder downsampling/middle blocks. Reusing more of those frozen computations is a concrete fallback if the small map cannot reproduce them. Model stitching gives precedent for learned interfaces between frozen networks, not a convergence guarantee for these codecs. Bansal et al., model stitching
Data and optimization
There are only 128 independent training prompt/seed trajectories per direction, yielding 896 correlated rows: six native clean estimates and one finished latent per trajectory. Validation has 16 prompts and 176 rows, including full-codec switched trajectories. At 180k steps, batch 16 means roughly 3,214 batch-equivalent passes over the same pool, not additional independent examples.
Training includes 94 photographic, 25 watercolor and nine oil-painting prompts. Charcoal and screen-print prompts occur in validation but not training. The worst forward prompt is a charcoal railway station, with prompt-mean LPIPS about 0.30. This suggests a coverage issue without proving style alone causes it. Switched validation cases are currently easier than native cases, particularly in reverse; their absence from training does not explain the current plateau.
Recent live minibatch latent NMSE is about 0.067 forward / 0.025 reverse, while EMA validation NMSE is about 0.252 / 0.060. These are not matched measurements: weights, batching and state mixtures differ. Native-only validation still has a substantial gap. A matched evaluation is required before concluding overfitting or deciding that a larger network is necessary.
Generator/encoder, critic and particle learning rates remain 2.4e-4, 3.6e-4 and 2.4e-3 respectively, with no refinement schedule. The forward critic readily distinguishes noisy residuals, but low discriminator loss is not proof of a saturated generator: for the implemented relativistic logistic loss, the generator score derivative approaches magnitude one at the observed positive real-minus-fake margin. Critic-gradient direction and propagation still may be poor. The score differences, rather than their absolute offset, drive the game. Relativistic GAN formulation
The critic noise stays at 1.3 after 8k. Lazy b-cap limits excessive gradient norms; it does not guarantee useful gradient directions. Every logged penalty falls on a lazy regularization step because 300 is divisible by four, so that graph is not an average per-update penalty. EMA's roughly 138-update half-life cannot alone explain a many-thousand-step plateau. Existing gradient/router diagnostics are largely restricted to step 1 and cannot diagnose the present state.
Ordered experiments and decision rules
These are conditional branches: a clear generalization gap moves data expansion ahead of capacity work. There is no requirement to run every candidate.
1. Establish a matched diagnosis
Evaluate the same selected EMA checkpoint on all native training and validation rows using identical decoding and metrics. Separate finished versus clean estimates, sigma/capture index, prompt, style/detail category, and switched validation. Separately compare live versus EMA weights from the same complete recovery snapshot; a selected best EMA export need not have matching saved live weights. Record source LPIPS, student-to-teacher LPIPS, pixel MSE, latent NMSE, per-prompt tails and images.
For gradient diagnostics, use one complete recovery snapshot containing its matching live generator and critic. Measure parameter and native-head-input gradient norms, adversarial versus latent-error gradient alignment, diagnostic perceptual-gradient alignment on a small subset, critic gradient quantiles, per-group update/weight ratios and particle GAN/VIC gradient ratios. Repeat across independent diagnostic noise draws. Preserve the training RNG. Diagnostic gradients do not add losses to training.
If training decoded error is low and validation error high, prioritize data. If both are high, test optimization and capacity. If gradients are noisy or systematically conflict with error reduction, investigate the critic before increasing its size. Also evaluate the same latent at nearby noise conditions: the teacher is independent of this condition once the latent is given.
2. Test a cheap refinement phase
Branch an unchanged-rate control and a quarter-rate candidate from the same full recovery snapshot per direction. Preserve all weights, optimizer moments, EMA, critic, particles, calibration, RNG, data, noise, paired-error Rp-logistic, b-cap and VIC. Give each 18k updates, validating every 3k. Candidate rates are G/encoder 6e-5, D 9e-5 and particles 6e-4.
Compare the final three validation milestones and selected checkpoints, rather than promoting a single transient minimum. Lower rates are a practical test, not a consequence of a theorem guaranteeing this Adam implementation converges. Separate GAN update rates have theoretical and empirical motivation under specific assumptions. TTUR
If useful, test a subsequent smaller-rate phase. If only training quality improves, move to data expansion. A D-only reduction is a later conditional arm if gradient diagnostics justify changing the game balance. Keep noise fixed while making that comparison. Instance-noise convergence results do not directly prove convergence of our shared-noise paired-error game with one-sided b-cap. Mescheder et al.
3. Separate objective limitations from representation limitations
Use 32–64 training-only latent examples for a diagnostic fit. Compare unchanged paired-error training with explicitly labeled supervised teacher-latent regression using the same architecture and examples. Success only with direct supervision implicates optimization/objective limitations; failure in both motivates a larger/restructured map or precision/gradient investigation. It does not by itself prove an information-theoretic impossibility.
If useful, test recipient-feature hints or teacher-distillation warmup as a separate recipe. Raw pre-normalization feature MSE can penalize differences removed by the native normalization; choose and justify the feature comparison. FitNets supports intermediate teacher guidance for small students, but applying it to these codec interfaces is an experimental inference. FitNets
Decoded pixel/perceptual terms are later explicit objective arms, evaluated with independent quality measures and images. They are not silently added to the existing ParticleGAN baseline. Perceptual supervision has precedent in trained feed-forward image transformations. Johnson et al.
4. Expand independent examples when the gap is confirmed
Start with 128→256 distinct training trajectories per direction, then 512 if needed. Include more styles, fine repeating detail, line art, surfaces, text on objects, low contrast and flat saturated regions, using objects, animals and landscapes. Hold out entire prompts and seeds. Add a larger independent development suite while retaining the old curves for continuity. Do not copy validation failures into training or open the reserved test.
The original pair preparation measured 1.2105 GPU-hours for 128 training and 16
validation cases, generating both directions. Observed training case intervals
averaged 27.23s. Adding 128 new training cases is approximately 1.0 GPU-hour including
loading; full 256/512/1024-case rebuilds with the same 16 validation cases are
roughly 2.2/4.1/8.0 GPU-hours. These are throughput extrapolations, not benchmarks
or completion promises; extra validation and resource contention add time.
Evidence: artifacts/runs/bridge-nova-qwen21-20260920/compute-ledger.json and
training-shard timestamps.
Multiple seeds for one prompt require the protocol validator to allow unique
(prompt, seed) pairs while retaining prompt-disjoint splits; it currently
rejects repeated prompt strings. 512 prompts × 2 seeds means 1,024 trajectories.
For a warm-start data-only experiment, retain the old normalization/calibration
and record the new data manifest explicitly. Fresh-seed qualification may then
calibrate on the new training data, identically within movable/fixed pairs.
Licensed external or synthetic images offer cheaper supplementary codec pairs: encode the source, then recompute the actual full-codec target from that latent. Encoding the original image independently into both codecs defines a different target. Native generated and intermediate latents remain necessary; encoded-image data do not automatically match their distribution. Apply image augmentation before encoding and regenerate targets; latent transforms need separate proof.
5. Test inexpensive capacity changes before a large model
Use one change per arm, first retaining the present objective and data. Compare both matched updates and measured GPU time. CPU architecture counts below exclude native prefix/head execution and the critic from convolution estimates.
| Candidate | Forward / reverse cost relative to current learned generator | Purpose |
|---|---|---|
| Zero-initialized raw-latent spatial bypass | About +5%/+7% convolution MACs; +104k/+154k parameters | Give detailed information a path around feature projection |
| Four→eight residual blocks, new blocks initialized as identities | About +32%/+28% convolution MACs; 2.882M/3.225M total trainable parameters | Test depth and learned spatial processing |
| Frozen recipient encoder middle block before current head | Benchmark actual latency/memory; reverse attention grid is larger | Reuse nonlinear native processing instead of relearning it |
| Width 96→192 | About 2.95×/2.64× convolution MACs; 6.756M/6.942M parameters | Later capacity test if smaller changes are inadequate |
The bypass packs raw inputs losslessly to the common 32² grid and adds raw→hidden and raw→recipient-feature paths. Zero weights preserve the original function. For extra residual blocks, zero the final convolution so each new block is an identity. Verify equality before training, migrate existing optimizer states by parameter name and initialize only newly introduced state. Function-preserving expansion is the relevant precedent, not permission to assume migration is automatically exact. Net2Net
Reusing the recipient middle block needs new feature calibration and native equality checks. If these variants fail, progressively move the source cut later and the recipient cut earlier. The complete decoder→encoder composition is the attainable quality reference at the expensive end of this quality/latency curve. The current compact size is a design choice, not a requirement to discard most pretrained computation.
6. Train and qualify the actual switching use case
After individual conversions improve, collect offline trajectories from frozen learned bridges on training-only prompts and predefined switch schedules. Label the encountered clean latents with the full-codec teacher, retain a fixed native mixture, and retrain. Current switched validation comes from full-codec rollouts, so it misses states induced by imperfect learned bridges. Dataset aggregation motivates this procedure; its imitation-learning guarantees do not automatically transfer to diffusion trajectories. DAgger
Evaluate finished conversion, round trips and early/late/repeated switches in both starting domains, with matched noise, guidance and denoiser-call budgets. Include full-codec controls and independent development prompts/seeds. All inference remains ordinary deterministic network forwards; teacher caches and rollout collection are training-only. 512px remains the qualified resolution until additional sizes are actually measured.
Promotion criteria and implementation requirements
A low or stationary GAN loss is not our convergence criterion. Each candidate must improve held-out decoded fidelity in the direction where it is adopted. Allow different winning recipes for forward and reverse, and qualify the combined pair. Retain source-LPIPS checkpoint selection for continuity; independently judge final-three-milestone student-to-teacher scores. Do not silently change the checkpoint-selection criterion. Proposed screening rules, to register before running the branches:
- Require at least 5% relative improvement in student-to-teacher LPIPS over its matched control, sustained in the mean of the final three milestones. Also report source LPIPS, pixel error and latent NMSE; teacher matching alone cannot establish source preservation.
- Reject a candidate with more than 5% relative deterioration in source LPIPS within a principal state-kind stratum unless explicitly treated as a separate task-specific tradeoff. Inspect worst-prompt tails and corresponding images.
- Bootstrap paired differences by prompt, keeping its seeds and trajectory rows together, or use a hierarchical prompt/seed bootstrap. Treat uncertainty from the current 16 prompts as exploratory; confirm promoted recipes on the larger predeclared development suite and independent seeds. Adaptive reuse can overfit a small validation set. Dwork et al.
- Stop extending an unchanged recipe when matched held-out quality remains flat across the registered evaluation window; diagnose or change one factor instead. A plateau alone is not distribution readiness.
- Before a repeatability or particle-advantage claim, run the selected recipe from at least three fresh seeds with matched movable/fixed clouds. Report every run and both directions. Reserve the existing test for the final comparison.
The 5% thresholds are practical proposed screening choices, not perceptual equivalence limits established by the literature. A distribution-quality claim also needs the wider use-case suite, visible artifact checks, exact export/native weight checks and measured inference latency/memory.
Optimization branches must use a complete recovery snapshot, not best EMA weights combined with an unrelated latest critic. Current continuation accepts only a larger horizon; LR, architecture, data and objective changes need a distinct recorded fork mechanism. Apply changed optimizer rates after loading saved state, verify resolved groups, and checkpoint schedules. Keep parent files immutable, record all changed fields, and preserve RNG and unchanged optimizer state.
Run one directional job per assigned GPU (forward 1, reverse 0), using the local venv and isolated Qwen runtime. Research did not interrupt the active 180k run or launch competing GPU work. The first implementation should be the matched audit and explicit fork support; the audit then determines which of data, optimization or architecture receives the next substantial compute allocation.