# Mini U-Net Colorizer: spatial artifacts and further training 20 September 2026. This study starts from the **bin-mapping repair**, not the broken Hub checkpoint. Improvements below must not be conflated with the previous repair's 55.5% reduction in chroma error. ## Selected release and final results **Selected:** mixed-data GPU checkpoint at update 748, from the completed 1,122-update run. Selection used 50 Imagenette plus 200 reserved COCO training images, with equal weight for each domain's relative chroma error. A candidate had to retain at least 90% of baseline predicted chroma in both domains and avoid worse fine excess-edge scores. CPU and GPU spatial-loss candidates failed the COCO color-retention gate (87.0% and 89.0%). The selected checkpoint retained 96.5% and 97.9%, respectively, with a 2.66% improvement in the selection score. No final-test results were used to change weights or decoding. All comparisons below start from the already repaired baseline. Lower error and excess-edge scores are better. Chroma ratio is predicted/reference mean chroma; it is not a realism score. | Test sample | Pipeline | Chroma error | Chroma ratio | Fine excess edge | Coarse excess edge | |---|---|---:|---:|---:|---:| | Imagenette 200 | Previous repair, raw | 13.318 | 0.965 | 0.5395 | 1.1962 | | Imagenette 200 | Previous repair, guided8 | 13.077 | 0.943 | 0.0829 | 0.5948 | | Imagenette 200 | New weights, raw | 13.056 | 0.919 | 0.5188 | 1.0782 | | Imagenette 200 | **New weights, guided8** | **12.847** | **0.899** | **0.0765** | **0.5354** | | COCO-val 100 | Previous repair, raw | 15.266 | 0.981 | 0.5896 | 1.4608 | | COCO-val 100 | Previous repair, guided8 | 14.992 | 0.959 | 0.0990 | 0.7479 | | COCO-val 100 | New weights, raw | 14.771 | 0.985 | 0.5813 | 1.3267 | | COCO-val 100 | **New weights, guided8** | **14.532** | **0.966** | **0.0923** | **0.6727** | Relative to the previous raw repair, the complete pipeline improves chroma error by **3.54% and 4.81%**, fine excess edges by **85.83% and 84.35%**, and coarse excess edges by **55.24% and 53.95%**, on Imagenette and COCO respectively. These are measured proxies, not percentages of visible blotches eliminated. Mean error improves on 142/200 and 76/100 images. Paired-image bootstrap 95% intervals for mean error reduction are [0.308, 0.639] and [0.362, 1.089] Lab units. There is a learned improvement beyond postprocessing: raw weights improve error by 1.97% and 3.25%; with the same guided8 decoder on both old and new weights, improvement is 1.76% and 3.07%. The corresponding guided-to-guided paired intervals are [0.060, 0.398] and [0.091, 0.812]. Per-image results, 10,000-resample bootstrap calculations and low-chroma-source summaries are provided in `reports/round2/final_results.json` and the final evaluation JSONs. Qualitative checks retain failures. The coffee image has substantially fewer fine rainbow patches but still contains wrong broad hues. The cat remains undercolored. The rocket's color error worsens from 21.94 to 22.86 despite smoothing. Four fixed probes and one fixed example per Imagenette class are shown in `final_visual.png` and `final_test_comparison.png`; they were not selected for favorable outcomes. The held-out quantitative pipeline resizes images to 256x256; the four application probes exercise the aspect-preserving, full-resolution output wrapper. Their metrics are not directly pooled. The result is a useful **app-testing candidate**, not evidence that all blotches are removed or that arbitrary real-world photos are production-ready. ## Deployment verification The selected safetensors file has SHA-256 `0e4c417375684a044860f8af3ac3a2fb44e1a5729254ca33abda758ea71aea6e`. All 65 learned parameter tensors changed from the repaired baseline; the verified color-bin buffer is identical. Parameter count remains 3,968,892. An ONNX opset-17 graph includes the learned model, temperature-0.38 decoding and guided8 filtering. It is 15,885,757 bytes and accepts dynamic spatial sizes and batches. Four numerical shape checks passed, with maximum Lab-ab difference below 0.000009 against PyTorch. Three image-wrapper comparisons preserved image dimensions and differed by at most one 8-bit RGB level. Ten code tests passed. The wrapper and graph input/output contract are documented in `DEPLOYMENT.md`. On this shared CPU host, ONNX Runtime 1.30.0 with two threads measured median 909ms and p90 1239ms for a 256x256 input (20 iterations after five warmups). Peak process RSS was 508MiB. This includes network, decoding and filtering, excludes file loading/full-resolution color conversion, and is not a browser or mobile latency guarantee. The graph is not quantized. Repository publication remains blocked: only the original Hugging Face connection is callable, with Jobs/read scopes and no write-repos scope. Previous direct and PR upload attempts returned 403; rechecking scopes at completion confirmed the same restriction. Neither main nor stable changed. The complete release, source, evidence and atomic main-upload helper are preserved for manual publication. No additional training is left running. ## Questions and evidence Three interventions were tested: spatial decoding, supervised spatial loss, and broader training data. Five training runs completed, including three L4 GPU runs. Every model retains 3,968,892 learned parameters and the same 236 color-bin meanings. Model selection uses validation images, with separate Imagenette and COCO external checks. Final quantitative results and the selected release follow below. ### Spatial decoding Thirteen alternatives were compared on 50 validation images: the repaired baseline; temperatures 0.6 and 0.8; logit pooling by factors 2, 4 and 8; luminance-guided filtering with radii 4, 8 and 16; horizontal-flip averaging; flip averaging plus radius-4 filtering; pooling plus filtering; and a gray control. The gray control achieves zero discontinuity scores, demonstrating why these scores cannot select a model by themselves. The [guided-filter paper](https://people.csail.mit.edu/kaiming/eccv10/index.html) motivates a local linear filter whose coefficients depend on luminance. This implementation filters predicted Lab chroma using only input luminance; reference colors are never passed into inference. It can smooth spurious color changes while retaining luminance boundaries. Same-luminance color boundaries remain a limitation. Radius 8 was nominated as the balanced single-pass mode using validation results and images. Radius 16 visibly removes more genuine small color details, despite better artifact scores. Flip averaging plus radius 4 is an optional two-pass mode. Temperature stays 0.38 and saturation stays 1.0. On the additional 200-image Imagenette test, radius-8 filtering alone changed mean chroma error from 13.318 to 13.077, fine excess color edge from 0.5395 to 0.0829, and coarse excess edge from 1.1962 to 0.5948. Mean predicted chroma relative to reference changed from 0.965 to 0.943. Error improved on 179/200 images; paired image bootstrap 95% interval for the mean improvement was [0.205, 0.276] Lab units. On 100 COCO-val photos, filtering alone changed error from 15.266 to 14.992, fine excess edge from 0.5896 to 0.0990, and coarse excess edge from 1.4608 to 0.7479. Chroma ratio changed from 0.981 to 0.959. Error improved on 95/100 images; the corresponding interval was [0.233, 0.314]. These intervals describe the image samples, not performance across arbitrary deployment photos. ### Learned spatial consistency The added loss matches chroma gradients to reference gradients at scales 1, 4 and 16. Chroma is normalized by 110; the gradient differences use smooth L1 with beta 0.05, averaged over scales and horizontal/vertical directions. Its coefficient is 10 alongside the original weighted classification loss. This is supervised transition matching, rather than a requirement that every region become uniformly gray. The CPU pair used 400 updates, batch 2, learning rate 2e-6 and seed 0. The GPU pair used 1,122 updates, batch 32, learning rate 1e-5 and seed 123: three complete passes through 11,943 training photos. Within each pair, the only training-objective difference is the spatial loss. Both use frozen BatchNorm statistics, identical data order, fixed audited color weights, and cosine learning-rate decay. CPU-to-GPU differences also change batch, precision, seed and training duration and are not a single-factor comparison. At the final CPU checkpoint, classification-only versus spatial loss produced validation error 12.384 versus 12.149, CE 2.8076 versus 2.7918, and chroma ratio 0.962 versus 0.923. At the final GPU checkpoint, error was 12.735 versus 12.326 and CE 2.8293 versus 2.8091. Earlier CPU checkpoints had lower error but appreciably less color, illustrating the need to inspect colorfulness. These results support the added loss in this experiment; they do not prove that the existing artifacts have one particular architectural cause. ### Broader data A third GPU arm added 5,382 eligible COCO training photos to the 11,943 Imagenette photos. The first two pinned COCO training shards contain 5,864 photos; 200 were reserved for validation and the remaining photos were filtered using the same low-chroma threshold. The combined training pool has 17,325 photos. The run used the classification objective and the same 1,122 updates and batch size as the GPU control, approximately 2.07 epochs over the larger pool. The original color vocabulary and class weights stayed fixed. The COCO training-validation set is distinct from the 100 COCO-val photos used as an external probe. Image-ID intersection between the training shards and that external sample is zero. The COCO revision is `26ddc382fe75dfc2a0655b5977e296ea10efebce`. The actual indices, hashes and training manifests are retained. Broader-data training was exploratory; this is not a claim that a small COCO subset is sufficient for production. ## What the diagnostics establish The constant-luminance and shifted-image probes show spatial sensitivity. Across ten validation images, shifting by one pixel changed predicted chroma by 2.31 Lab units on average in the central region; an eight-pixel shift changed it by 1.45. Flat-input responses also contain spatial color variation. This does not isolate a causal layer: padding, pooling, dilations, upsampling and learned features can all contribute. [Odena et al.](https://distill.pub/2016/deconv-checkerboard/) motivates testing upsampling effects, but the current kernel-2/stride-2 layers do not have the classic uneven-overlap configuration. Replacing them without a controlled training comparison is not established here as a fix. The current context path already covers hundreds of input pixels; simply increasing dilation is not a demonstrated remedy. The remaining wrong hues also point to a semantic problem that spatial filtering cannot solve. The [original colorization research](https://richzhang.github.io/colorization/) discusses dataset bias despite much larger training coverage, and [DDColor](https://arxiv.org/abs/2212.11613) motivates semantic and multiscale features. Semantic teacher-to-student training or a pretrained compact encoder remain untested architectural alternatives here, not completed work. ## Evaluation limits The 200-image Imagenette test is disjoint from the previous 100-image test and the 50-image validation set. It comes from the historical seed-0 5% holdout. Exposure during older upstream training remains unknown. The 100 COCO-val images are the first viewer rows, a convenience sample with 72 object categories represented, not a population-representative random test. Chroma error measures closeness to one reference colorization; plausible alternative colors may score poorly. Excess-edge measures are heuristics, not counts of visible blotches or human quality ratings. Grayscale-source references are separately counted and color-source-only summaries retained. Qualitative examples include known failures, not only favorable results. The additional 200-photo COCO training-validation comparison uses the reserved indices, with 256px JPEG95 copies for transfer. Every candidate receives the same decoded copies. These images are used for selecting weights and are not reported as an independent final test. ## Execution and reproducibility The custom plugin named Huggingface still exposed no tools. The original Hugging Face connection supplied Jobs and read access but lacked write-repos. A 17 MiB checksum-verified log roundtrip established a fallback artifact transport. Larger two-checkpoint archives exceeded the tool's 64 MiB response frame because the response repeats log content; CPU helpers exported one checkpoint at a time. Explicit log-tail limits avoided server truncation. All recovered archives are checked by size, SHA-256 and ZIP CRC. Two initial GPU starts stopped on nonfinite FP16 gradients before completing training. BF16 restarts completed with finite-gradient guards enabled. Those failed starts are recorded separately from the five completed training runs. The three GPU jobs were `6aaf8ceb52d0dbd7f1d73094`, `6aaf8cf752d0dbd7f1d7309a`, and `6aaf8dad51992417dfccc557`. Both best and last learned checkpoints are retained. GPU optimizer state was not exported; those checkpoints support a new-optimizer fine-tune, not exact training resume. CPU optimizer states are retained. Source, configurations, environments, priors, split definitions and per-image results accompany them. The corrected decoder and vocabulary guards continue to prevent the earlier silent mapping mismatch.