File size: 4,259 Bytes
1cca2dd | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 | # Final training decision — 29 September 2026
All seven pilot arms completed. Reviewed all 32 fixed COCO examples in each of the single, multi and encoder comparison groups, all14 historical photos in production-pipeline grids, and the14-photo multi-palette archive comparison. This is qualitative assistant review, not an independent human preference study.
## Decision
Continue the MobileNetV3 four-palette model with mixed complete-image teacher/original targets, initialized from `experiments/decision-20260928-v2/multi_mix/candidate` (pilot step2000). Complete deployed model:3,995,832 parameters. This is a modest best-compromise candidate, not a proven major advance over v3. Retain released v3 until final visual review.
## Evidence
Relative to v3, multi_mix guided patch excess decreased4.69% on ordinary grayscale (0.81598→0.77772) and9.37% on film grayscale (0.77398→0.70148). Exploratory paired95% intervals for the absolute changes were[-0.07208,-0.00611] and[-0.10446,-0.04150]. These intervals do not correct for selecting among seven experiments.
Mean chroma was almost retained:14.0299→13.9184, and13.7452→13.6813. Coverage nevertheless fell from54.44%→52.63% and52.99%→52.30%; missed-colour rates slightly improved, with uncertain paired differences. Original-colour error changed little. No metric is a direct perceptual-quality guarantee.
MobileViT had the lowest patch metric but its images often had a yellow/brown cast: skateboarder background, tennis player and court, lake, hand holding food, kitchen hand, and indoor scenes. It also lost useful colour distinctions in the dance scene and food images. It improved some landscape/elephant details and the historical children/sky, but the overall trade-off is unsuitable for release. Its selected step1000 precedes its staged curriculum, so those scores do not validate that curriculum.
Single mixed/staged outputs were generally close to v3 and sometimes smoother, but reduced colour coverage and saturation. Single_staged selected step1000, before the late robustness phase; it is not evidence that the late film/consistency phase worked. Multi_staged retained more colour but its ordinary-grayscale patch score worsened. Thus the final course uses mixed targets without the unvalidated late-stage bundle.
Multi_mix preserves recognisable blue sky, green foliage and skin distinctions in the modern examples, with less evidence of the broad sepia substitution seen in MobileViT. Historical images remain substantially similar to v3. The old man's clothes and the children's field remain muted; the powerhouse pipe still changes hue along its length. The power station still receives a pink cast. No claim that blotches are eliminated.
The multi selector chose modes[39,8,1,0] on48 diagnostic images; alternatives are unevenly used. Default target loss0.5174 versus oracle0.4234 shows room for selection improvement, but oracle performance is not the deployable result. Resize changed the selected mode in1/48 examples. Keep all parameters and the learned default; do not choose an output using unknown ground truth at inference.
## Final run
Up to24,000 additional updates, batch24, all16,230 existing filtered training images, cached teacher/original targets sampled50/50 per image. Fresh optimizer; encoder LR2e-6 and decoder LR2e-5, warmup150, cosine floor15%, frozen encoder batch statistics, clipping1. No RL, critic, new teacher generation, or architecture expansion. Save every2000 updates and retain the initial pilot as a valid candidate. Stop after six evaluations without a development-score improvement of0.002. The development score retains colour/error guardrails.
New development/test samples exclude the pilot's development/test IDs and the earlier evaluation IDs. Test opens after development checkpoint selection. Final outputs and image comparisons are pushed under `experiments/final-20260929` on main; root production weights are not automatically overwritten. A lower metric is insufficient for promotion.
Budget: one L4,60min hard timeout ($0.80 maximum at the currently documented rate),50min internal training limit. Expected about25–40min from measured pilot throughput, with setup and evaluation overhead. No automatic retry jobs.
|