# Final checkpoint review — 29 September 2026 ## Decision Keep **v3.0** as the production model. The completed four-palette multi_mix continuation is retained on main at experiments/final-20260929/multi_mix/candidate, but is not promoted to root weights or the Space default. The stable branch is unchanged. The final run completed 16,000 additional updates, then stopped by the prespecified development patience rule. The development-selected candidate was step 4,000 (after the prior 2,000-step pilot), not the last training step. It has 3,995,832 parameters in total including the MobileNetV3 encoder. ## Held-out comparison against v3 | Measure | Gray v3 → candidate | Film gray v3 → candidate | |---|---:|---:| | Guided patch excess (lower) | 0.794 → 0.767 | 0.745 → 0.698 | | Raw patch excess (lower) | 0.996 → 0.947 | 0.937 → 0.862 | | Neutral spill (lower) | 0.234 → 0.197 | 0.245 → 0.216 | | Colour coverage (higher) | 51.50% → 48.38% | 50.68% → 47.98% | | Mean chroma | 13.24 → 12.54 | 12.95 → 12.33 | | Lab ab error (lower) | 12.64 → 12.52 | 12.79 → 12.71 | Exploratory paired 95% intervals for guided patch change are [-0.0619, +0.0049] on gray and [-0.0800, -0.0145] on film. The corresponding coverage drops have intervals [-0.0468, -0.0158] and [-0.0437, -0.0108]. This is not a human preference study. I reviewed all fourteen archival comparisons, held-out modern production grids and special probes. Most archival outputs are nearly identical and remain warm/sepia or muted. The candidate smooths some local artifacts and preserves useful sky/foliage in a few images, but more clothes, interiors and backgrounds remain essentially gray. The Moon probe is still implausibly pink. Neither model eliminates residual blotches. The modest patch improvement does not compensate for the consistent loss of colour on the user's main grayscale-photo use case. The pilot had a more favourable colour-retention trade-off; this continuation moved toward desaturation. Selection within the final course did not establish that it surpasses the released v3. Checkpoints, metrics, source, diagnostics and comparison grids are saved under experiments/final-20260929, with summary.json and multi_mix_test_paired.json providing exact figures. ## Production state Root model.safetensors, config.json and colorizer.onnx remain v3.0, with 3,994,676 total learned parameters. The ZeroGPU Space remains pinned to immutable v3 revision 1a9eb8af2754ad2329a24cfe50d388cb559441d0 and live inference was verified on a historical photograph. No further training job is scheduled. This is the best release among the tested candidates, not evidence that no future sub-4M model could improve. ## Complete image inspection and palette check The final run contains 27 JPEG comparison grids. All 27 were inspected: four modern multi-test grids, four single-baseline grids, four production-test grids, fourteen archive comparisons across their single/multi/production layouts, both probe panels, and five v3-only resolution grids. The single-baseline and resolution grids establish context rather than additional candidate wins. The v3 384/512 examples show that a detail-setting change can alter the colour choice, especially for historical scenes, but increasing resolution is not a universal improvement. The multi-model examples are usually extremely close to v3. The candidate sometimes smooths a colour edge or tones down an implausible patch, including the Moon's localized red mark. It still paints the Moon broadly pink. In the historical set, some faces and greenery are coherent; hats, clothing, interiors and the power-station surfaces often remain neutral or acquire a global brown/purple cast. Several ordinary outdoor samples keep recognizable blue sky and green foliage. This is a modest trade-off, not a decisive visual quality gain. An additional six-archive-photo check rendered **all four modes**, using the saved selected checkpoint and the production guided-upsample renderer. The [comparison panel](FINAL_PALETTE_REVIEW.jpg) and [mode probabilities](FINAL_PALETTE_REVIEW.json) are saved here. Each image selects mode 0 by default, with its probability only about 0.30–0.34 and the other three typically about 0.18–0.27. Modes primarily vary the global tint: reddish/purple, warm sepia, near-neutral, and cool/bluish. They do not create convincing object-level alternative interpretations. On the power station, mode 2 removes the purple tint but largely returns the structure to gray; on the children-and-beets scene, the warm mode does not discover a strongly coloured field. Thus the extra palette control is not a reason to promote the checkpoint. This assessment is a visual comparison of selected examples and an engineering diagnostic, not a blinded human study of historical plausibility. The held-out 160-image paired metrics remain the quantitative check.