File size: 18,721 Bytes
1f167a1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
# Nova/Qwen convergence strategy, September 21

Three independent research reviews covered capacity/pretrained reuse,
optimization, and interface/data/evaluation limits. The synthesis below combines
their code audits, primary literature, live metrics and existing validation
images. This is a proposed experiment sequence; it does not modify the running
180k campaign or its matched controls.

The best-supported hypothesis is a combination of limited training coverage and
a restrictive learned approximation, with constant-rate optimization adding
noise. There is no demonstrated codec incompatibility preventing a much more
accurate converter. No experiment or cited theorem guarantees that the present
small adversarially trained network will reach the full-codec reference.

## Evidence and what it establishes

The accompanying [measurement snapshot](nova-qwen21-convergence-evidence-20260921.json)
records the actual run paths, read time, quality strata and training diagnostics.
At that snapshot:

| Direction | Selected step | Source LPIPS | Student-to-teacher LPIPS | Full-codec source LPIPS |
|---|---:|---:|---:|---:|
| Nova → Qwen | 126,000 | 0.138064 | 0.137344 | 0.010601 |
| Qwen → Nova | 120,000 | 0.119970 | 0.098619 | 0.054272 |

Forward's median validation LPIPS fell from 0.18963 over 36k–60k to 0.14715 over
93k–120k. Reverse's corresponding medians were 0.12425 and 0.12037; over
123k–150k its median was 0.12026. Forward still improves, while reverse is nearly
flat. A finished forward kingfisher at 126k preserves composition but has clear
block and edge distortions absent from the source and full-codec target.

The exact teacher is
`T(z) = E_target(quantize(channel_convert(clamp(D_source(z)))))`, using posterior
mode. It needs only the raw latent already supplied to the learned bridge.
The native decoder remainder consumes the retained source-prefix features, so
those features are sufficient to compute the teacher in principle. Teacher
targets also pass through the recipient head that the learned bridge uses.
The restriction is the compact map inserted between the frozen endpoints.
Native prefix/head, stored-pair, cache and exported-forward checks have passed;
see [the original diagnosis](nova-qwen21-forward-repair-20260921.md).

The full-codec score is a measured reference, not a mathematical lower bound.
Student-to-teacher fidelity and source preservation must both be measured.
General exact reversibility is impossible across the declared RGBA→RGB
compositing operation: alpha is discarded. Quantization and codec compression
also limit round trips. Equal-dimensional Gaussian noise transport does not
align the two diffusion vector fields, so conversion accuracy and switched
generation require separate qualification.

## Capacity and pretrained processing

| Property | Forward | Reverse |
|---|---:|---:|
| Raw input | 16×64×64 | 64×32×32 |
| Number of latent scalars | 65,536 | 65,536 |
| Native feature channels before learned projection | 384 | 1,152 |
| Projected channels | 96 | 96 |
| Learned residual blocks | 4 | 4 |
| Predicted recipient feature channels | 768 | 384 |
| Trainable parameters | 2,217,828 | 2,560,356 |

Pixel shuffle/unshuffle is lossless. Detailed spatial information instead passes
through a pointwise 96-channel projection, without an independent raw-latent
spatial bypass. This limits the function class but does not prove information
loss on the actual data manifold. The source prefix already contains native
attention; the whole bridge does not lack global context. The learned global
particle condition is only four-dimensional.

The learned map skips source decoder upsampling blocks and recipient encoder
downsampling/middle blocks. Reusing more of those frozen computations is a
concrete fallback if the small map cannot reproduce them. Model stitching gives
precedent for learned interfaces between frozen networks, not a convergence
guarantee for these codecs. [Bansal et al., model stitching](https://arxiv.org/abs/2106.07682)

## Data and optimization

There are only 128 independent training prompt/seed trajectories per direction,
yielding 896 correlated rows: six native clean estimates and one finished latent
per trajectory. Validation has 16 prompts and 176 rows, including full-codec
switched trajectories. At 180k steps, batch 16 means roughly 3,214 batch-equivalent
passes over the same pool, not additional independent examples.

Training includes 94 photographic, 25 watercolor and nine oil-painting prompts.
Charcoal and screen-print prompts occur in validation but not training. The
worst forward prompt is a charcoal railway station, with prompt-mean LPIPS about
0.30. This suggests a coverage issue without proving style alone causes it.
Switched validation cases are currently easier than native cases, particularly
in reverse; their absence from training does not explain the current plateau.

Recent live minibatch latent NMSE is about 0.067 forward / 0.025 reverse, while
EMA validation NMSE is about 0.252 / 0.060. These are not matched measurements:
weights, batching and state mixtures differ. Native-only validation still has a
substantial gap. A matched evaluation is required before concluding overfitting
or deciding that a larger network is necessary.

Generator/encoder, critic and particle learning rates remain 2.4e-4, 3.6e-4 and
2.4e-3 respectively, with no refinement schedule. The forward critic readily
distinguishes noisy residuals, but low discriminator loss is not proof of a
saturated generator: for the implemented relativistic logistic loss, the
generator score derivative approaches magnitude one at the observed positive
real-minus-fake margin. Critic-gradient direction and propagation still may be
poor. The score differences, rather than their absolute offset, drive the game.
[Relativistic GAN formulation](https://arxiv.org/abs/1807.00734)

The critic noise stays at 1.3 after 8k. Lazy b-cap limits excessive gradient norms;
it does not guarantee useful gradient directions. Every logged penalty falls
on a lazy regularization step because 300 is divisible by four, so that graph
is not an average per-update penalty. EMA's roughly 138-update half-life cannot
alone explain a many-thousand-step plateau. Existing gradient/router diagnostics
are largely restricted to step 1 and cannot diagnose the present state.

## Ordered experiments and decision rules

These are conditional branches: a clear generalization gap moves data expansion
ahead of capacity work. There is no requirement to run every candidate.

### 1. Establish a matched diagnosis

Evaluate the same selected EMA checkpoint on all native training and validation
rows using identical decoding and metrics. Separate finished versus clean
estimates, sigma/capture index, prompt, style/detail category, and switched
validation. Separately compare live versus EMA weights from the same complete
recovery snapshot; a selected best EMA export need not have matching saved live
weights. Record source LPIPS,
student-to-teacher LPIPS, pixel MSE, latent NMSE, per-prompt tails and images.

For gradient diagnostics, use one complete recovery snapshot containing its
matching live generator and critic. Measure parameter and native-head-input
gradient norms, adversarial versus latent-error gradient alignment, diagnostic
perceptual-gradient alignment on a small subset, critic gradient quantiles,
per-group update/weight ratios and particle GAN/VIC gradient ratios. Repeat
across independent diagnostic noise draws. Preserve the training RNG. Diagnostic
gradients do not add losses to training.

If training decoded error is low and validation error high, prioritize data.
If both are high, test optimization and capacity. If gradients are noisy or
systematically conflict with error reduction, investigate the critic before
increasing its size. Also evaluate the same latent at nearby noise conditions:
the teacher is independent of this condition once the latent is given.

### 2. Test a cheap refinement phase

Branch an unchanged-rate control and a quarter-rate candidate from the same
full recovery snapshot per direction. Preserve all weights, optimizer moments,
EMA, critic, particles, calibration, RNG, data, noise, paired-error Rp-logistic,
b-cap and VIC. Give each 18k updates, validating every 3k. Candidate rates are
G/encoder 6e-5, D 9e-5 and particles 6e-4.

Compare the final three validation milestones and selected checkpoints, rather
than promoting a single transient minimum. Lower rates are a practical test,
not a consequence of a theorem guaranteeing this Adam implementation converges.
Separate GAN update rates have theoretical and empirical motivation under
specific assumptions. [TTUR](https://arxiv.org/abs/1706.08500)

If useful, test a subsequent smaller-rate phase. If only training quality
improves, move to data expansion. A D-only reduction is a later conditional arm
if gradient diagnostics justify changing the game balance. Keep noise fixed
while making that comparison. Instance-noise convergence results do not directly
prove convergence of our shared-noise paired-error game with one-sided b-cap.
[Mescheder et al.](https://proceedings.mlr.press/v80/mescheder18a.html)

### 3. Separate objective limitations from representation limitations

Use 32–64 training-only latent examples for a diagnostic fit. Compare unchanged
paired-error training with explicitly labeled supervised teacher-latent
regression using the same architecture and examples. Success only with direct
supervision implicates optimization/objective limitations; failure in both
motivates a larger/restructured map or precision/gradient investigation. It does
not by itself prove an information-theoretic impossibility.

If useful, test recipient-feature hints or teacher-distillation warmup as a
separate recipe. Raw pre-normalization feature MSE can penalize differences
removed by the native normalization; choose and justify the feature comparison.
FitNets supports intermediate teacher guidance for small students, but applying
it to these codec interfaces is an experimental inference. [FitNets](https://arxiv.org/abs/1412.6550)

Decoded pixel/perceptual terms are later explicit objective arms, evaluated with
independent quality measures and images. They are not silently added to the
existing ParticleGAN baseline. Perceptual supervision has precedent in trained
feed-forward image transformations. [Johnson et al.](https://cs.stanford.edu/people/jcjohns/eccv16/)

### 4. Expand independent examples when the gap is confirmed

Start with 128→256 distinct training trajectories per direction, then 512 if
needed. Include more styles, fine repeating detail, line art, surfaces, text on
objects, low contrast and flat saturated regions, using objects, animals and
landscapes. Hold out entire prompts and seeds. Add a larger independent
development suite while retaining the old curves for continuity. Do not copy
validation failures into training or open the reserved test.

The original pair preparation measured 1.2105 GPU-hours for 128 training and 16
validation cases, generating both directions. Observed training case intervals
averaged 27.23s. Adding 128 new training cases is approximately 1.0 GPU-hour including
loading; full 256/512/1024-case rebuilds with the same 16 validation cases are
roughly 2.2/4.1/8.0 GPU-hours. These are throughput extrapolations, not benchmarks
or completion promises; extra validation and resource contention add time.
Evidence: `artifacts/runs/bridge-nova-qwen21-20260920/compute-ledger.json` and
training-shard timestamps.

Multiple seeds for one prompt require the protocol validator to allow unique
`(prompt, seed)` pairs while retaining prompt-disjoint splits; it currently
rejects repeated prompt strings. 512 prompts × 2 seeds means 1,024 trajectories.
For a warm-start data-only experiment, retain the old normalization/calibration
and record the new data manifest explicitly. Fresh-seed qualification may then
calibrate on the new training data, identically within movable/fixed pairs.

Licensed external or synthetic images offer cheaper supplementary codec pairs:
encode the source, then recompute the actual full-codec target from that latent.
Encoding the original image independently into both codecs defines a different
target. Native generated and intermediate latents remain necessary; encoded-image
data do not automatically match their distribution. Apply image augmentation
before encoding and regenerate targets; latent transforms need separate proof.

### 5. Test inexpensive capacity changes before a large model

Use one change per arm, first retaining the present objective and data. Compare
both matched updates and measured GPU time. CPU architecture counts below
exclude native prefix/head execution and the critic from convolution estimates.

| Candidate | Forward / reverse cost relative to current learned generator | Purpose |
|---|---:|---|
| Zero-initialized raw-latent spatial bypass | About +5%/+7% convolution MACs; +104k/+154k parameters | Give detailed information a path around feature projection |
| Four→eight residual blocks, new blocks initialized as identities | About +32%/+28% convolution MACs; 2.882M/3.225M total trainable parameters | Test depth and learned spatial processing |
| Frozen recipient encoder middle block before current head | Benchmark actual latency/memory; reverse attention grid is larger | Reuse nonlinear native processing instead of relearning it |
| Width 96→192 | About 2.95×/2.64× convolution MACs; 6.756M/6.942M parameters | Later capacity test if smaller changes are inadequate |

The bypass packs raw inputs losslessly to the common 32² grid and adds raw→hidden
and raw→recipient-feature paths. Zero weights preserve the original function.
For extra residual blocks, zero the final convolution so each new block is an
identity. Verify equality before training, migrate existing optimizer states by
parameter name and initialize only newly introduced state. Function-preserving
expansion is the relevant precedent, not permission to assume migration is
automatically exact. [Net2Net](https://arxiv.org/abs/1511.05641)

Reusing the recipient middle block needs new feature calibration and native
equality checks. If these variants fail, progressively move the source cut later
and the recipient cut earlier. The complete decoder→encoder composition is the
attainable quality reference at the expensive end of this quality/latency curve.
The current compact size is a design choice, not a requirement to discard most
pretrained computation.

### 6. Train and qualify the actual switching use case

After individual conversions improve, collect offline trajectories from frozen
learned bridges on training-only prompts and predefined switch schedules. Label
the encountered clean latents with the full-codec teacher, retain a fixed native
mixture, and retrain. Current switched validation comes from full-codec rollouts,
so it misses states induced by imperfect learned bridges. Dataset aggregation
motivates this procedure; its imitation-learning guarantees do not automatically
transfer to diffusion trajectories. [DAgger](https://proceedings.mlr.press/v15/ross11a.html)

Evaluate finished conversion, round trips and early/late/repeated switches in
both starting domains, with matched noise, guidance and denoiser-call budgets.
Include full-codec controls and independent development prompts/seeds. All
inference remains ordinary deterministic network forwards; teacher caches and
rollout collection are training-only. 512px remains the qualified resolution
until additional sizes are actually measured.

## Promotion criteria and implementation requirements

A low or stationary GAN loss is not our convergence criterion. Each candidate
must improve held-out decoded fidelity in the direction where it is adopted.
Allow different winning recipes for forward and reverse, and qualify the combined
pair. Retain source-LPIPS checkpoint selection for continuity; independently
judge final-three-milestone student-to-teacher scores. Do not silently change the
checkpoint-selection criterion. Proposed screening rules, to register before
running the branches:

- Require at least 5% relative improvement in student-to-teacher LPIPS over its
  matched control, sustained in the mean of the final three milestones. Also
  report source LPIPS, pixel error and latent NMSE; teacher matching alone cannot
  establish source preservation.
- Reject a candidate with more than 5% relative deterioration in source LPIPS
  within a principal state-kind stratum unless explicitly treated as a separate
  task-specific tradeoff. Inspect worst-prompt tails and corresponding images.
- Bootstrap paired differences by prompt, keeping its seeds and trajectory rows
  together, or use a hierarchical prompt/seed bootstrap. Treat uncertainty from the current 16 prompts as exploratory;
  confirm promoted recipes on the larger predeclared development suite and
  independent seeds. Adaptive reuse can overfit a small validation set.
  [Dwork et al.](https://arxiv.org/abs/1506.02629)
- Stop extending an unchanged recipe when matched held-out quality remains flat
  across the registered evaluation window; diagnose or change one factor instead.
  A plateau alone is not distribution readiness.
- Before a repeatability or particle-advantage claim, run the selected recipe
  from at least three fresh seeds with matched movable/fixed clouds. Report every
  run and both directions. Reserve the existing test for the final comparison.

The 5% thresholds are practical proposed screening choices, not perceptual
equivalence limits established by the literature. A distribution-quality claim
also needs the wider use-case suite, visible artifact checks, exact export/native
weight checks and measured inference latency/memory.

Optimization branches must use a complete recovery snapshot, not best EMA weights
combined with an unrelated latest critic. Current continuation accepts only a
larger horizon; LR, architecture, data and objective changes need a distinct
recorded fork mechanism. Apply changed optimizer rates after loading saved state,
verify resolved groups, and checkpoint schedules. Keep parent files immutable,
record all changed fields, and preserve RNG and unchanged optimizer state.

Run one directional job per assigned GPU (forward 1, reverse 0), using the local
venv and isolated Qwen runtime. Research did not interrupt the active 180k run or
launch competing GPU work. The first implementation should be the matched audit
and explicit fork support; the audit then determines which of data, optimization
or architecture receives the next substantial compute allocation.