File size: 13,661 Bytes
704fa80 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 | # Mini U-Net Colorizer: spatial artifacts and further training
20 September 2026. This study starts from the **bin-mapping repair**, not the
broken Hub checkpoint. Improvements below must not be conflated with the
previous repair's 55.5% reduction in chroma error.
## Selected release and final results
**Selected:** mixed-data GPU checkpoint at update 748, from the completed
1,122-update run. Selection used 50 Imagenette plus 200 reserved COCO training
images, with equal weight for each domain's relative chroma error. A candidate
had to retain at least 90% of baseline predicted chroma in both domains and
avoid worse fine excess-edge scores. CPU and GPU spatial-loss candidates
failed the COCO color-retention gate (87.0% and 89.0%). The selected checkpoint
retained 96.5% and 97.9%, respectively, with a 2.66% improvement in the
selection score. No final-test results were used to change weights or decoding.
All comparisons below start from the already repaired baseline. Lower error
and excess-edge scores are better. Chroma ratio is predicted/reference mean
chroma; it is not a realism score.
| Test sample | Pipeline | Chroma error | Chroma ratio | Fine excess edge | Coarse excess edge |
|---|---|---:|---:|---:|---:|
| Imagenette 200 | Previous repair, raw | 13.318 | 0.965 | 0.5395 | 1.1962 |
| Imagenette 200 | Previous repair, guided8 | 13.077 | 0.943 | 0.0829 | 0.5948 |
| Imagenette 200 | New weights, raw | 13.056 | 0.919 | 0.5188 | 1.0782 |
| Imagenette 200 | **New weights, guided8** | **12.847** | **0.899** | **0.0765** | **0.5354** |
| COCO-val 100 | Previous repair, raw | 15.266 | 0.981 | 0.5896 | 1.4608 |
| COCO-val 100 | Previous repair, guided8 | 14.992 | 0.959 | 0.0990 | 0.7479 |
| COCO-val 100 | New weights, raw | 14.771 | 0.985 | 0.5813 | 1.3267 |
| COCO-val 100 | **New weights, guided8** | **14.532** | **0.966** | **0.0923** | **0.6727** |
Relative to the previous raw repair, the complete pipeline improves chroma
error by **3.54% and 4.81%**, fine excess edges by **85.83% and 84.35%**, and
coarse excess edges by **55.24% and 53.95%**, on Imagenette and COCO respectively.
These are measured proxies, not percentages of visible blotches eliminated.
Mean error improves on 142/200 and 76/100 images. Paired-image bootstrap 95%
intervals for mean error reduction are [0.308, 0.639] and [0.362, 1.089] Lab units.
There is a learned improvement beyond postprocessing: raw weights improve
error by 1.97% and 3.25%; with the same guided8 decoder on both old and new
weights, improvement is 1.76% and 3.07%. The corresponding guided-to-guided
paired intervals are [0.060, 0.398] and [0.091, 0.812]. Per-image results,
10,000-resample bootstrap calculations and low-chroma-source summaries are
provided in `reports/round2/final_results.json` and the final evaluation JSONs.
Qualitative checks retain failures. The coffee image has substantially fewer
fine rainbow patches but still contains wrong broad hues. The cat remains
undercolored. The rocket's color error worsens from 21.94 to 22.86 despite
smoothing. Four fixed probes and one fixed example per Imagenette class are
shown in `final_visual.png` and `final_test_comparison.png`; they were not
selected for favorable outcomes. The held-out quantitative pipeline resizes
images to 256x256; the four application probes exercise the aspect-preserving,
full-resolution output wrapper. Their metrics are not directly pooled.
The result is a useful **app-testing candidate**, not evidence that all
blotches are removed or that arbitrary real-world photos are production-ready.
## Deployment verification
The selected safetensors file has SHA-256
`0e4c417375684a044860f8af3ac3a2fb44e1a5729254ca33abda758ea71aea6e`.
All 65 learned parameter tensors changed from the repaired baseline; the
verified color-bin buffer is identical. Parameter count remains 3,968,892.
An ONNX opset-17 graph includes the learned model, temperature-0.38 decoding
and guided8 filtering. It is 15,885,757 bytes and accepts dynamic spatial sizes
and batches. Four numerical shape checks passed, with maximum Lab-ab difference
below 0.000009 against PyTorch. Three image-wrapper comparisons preserved image
dimensions and differed by at most one 8-bit RGB level. Ten code tests passed.
The wrapper and graph input/output contract are documented in `DEPLOYMENT.md`.
On this shared CPU host, ONNX Runtime 1.30.0 with two threads measured median
909ms and p90 1239ms for a 256x256 input (20 iterations after five warmups).
Peak process RSS was 508MiB. This includes network, decoding and filtering,
excludes file loading/full-resolution color conversion, and is not a browser
or mobile latency guarantee. The graph is not quantized.
Repository publication remains blocked: only the original Hugging Face
connection is callable, with Jobs/read scopes and no write-repos scope.
Previous direct and PR upload attempts returned 403; rechecking scopes at
completion confirmed the same restriction. Neither main nor stable changed.
The complete release, source, evidence and atomic main-upload helper are
preserved for manual publication. No additional training is left running.
## Questions and evidence
Three interventions were tested: spatial decoding, supervised spatial loss,
and broader training data. Five training runs completed, including three L4
GPU runs. Every model retains 3,968,892 learned parameters and the same 236
color-bin meanings. Model selection uses validation images, with separate
Imagenette and COCO external checks. Final quantitative results and the selected release follow below.
### Spatial decoding
Thirteen alternatives were compared on 50 validation images: the repaired
baseline; temperatures 0.6 and 0.8; logit pooling by factors 2, 4 and 8;
luminance-guided filtering with radii 4, 8 and 16; horizontal-flip averaging;
flip averaging plus radius-4 filtering; pooling plus filtering; and a gray
control. The gray control achieves zero discontinuity scores, demonstrating
why these scores cannot select a model by themselves.
The [guided-filter paper](https://people.csail.mit.edu/kaiming/eccv10/index.html)
motivates a local linear filter whose coefficients depend on luminance. This
implementation filters predicted Lab chroma using only input luminance;
reference colors are never passed into inference. It can smooth spurious
color changes while retaining luminance boundaries. Same-luminance color
boundaries remain a limitation.
Radius 8 was nominated as the balanced single-pass mode using validation
results and images. Radius 16 visibly removes more genuine small color
details, despite better artifact scores. Flip averaging plus radius 4 is an
optional two-pass mode. Temperature stays 0.38 and saturation stays 1.0.
On the additional 200-image Imagenette test, radius-8 filtering alone changed
mean chroma error from 13.318 to 13.077, fine excess color edge from 0.5395 to
0.0829, and coarse excess edge from 1.1962 to 0.5948. Mean predicted chroma
relative to reference changed from 0.965 to 0.943. Error improved on 179/200
images; paired image bootstrap 95% interval for the mean improvement was
[0.205, 0.276] Lab units.
On 100 COCO-val photos, filtering alone changed error from 15.266 to 14.992,
fine excess edge from 0.5896 to 0.0990, and coarse excess edge from 1.4608 to
0.7479. Chroma ratio changed from 0.981 to 0.959. Error improved on 95/100
images; the corresponding interval was [0.233, 0.314]. These intervals describe
the image samples, not performance across arbitrary deployment photos.
### Learned spatial consistency
The added loss matches chroma gradients to reference gradients at scales
1, 4 and 16. Chroma is normalized by 110; the gradient differences use smooth
L1 with beta 0.05, averaged over scales and horizontal/vertical directions.
Its coefficient is 10 alongside the original weighted classification loss.
This is supervised transition matching, rather than a requirement that every
region become uniformly gray.
The CPU pair used 400 updates, batch 2, learning rate 2e-6 and seed 0. The
GPU pair used 1,122 updates, batch 32, learning rate 1e-5 and seed 123: three
complete passes through 11,943 training photos. Within each pair, the only
training-objective difference is the spatial loss. Both use frozen BatchNorm
statistics, identical data order, fixed audited color weights, and cosine
learning-rate decay. CPU-to-GPU differences also change batch, precision,
seed and training duration and are not a single-factor comparison.
At the final CPU checkpoint, classification-only versus spatial loss produced
validation error 12.384 versus 12.149, CE 2.8076 versus 2.7918, and chroma
ratio 0.962 versus 0.923. At the final GPU checkpoint, error was 12.735 versus
12.326 and CE 2.8293 versus 2.8091. Earlier CPU checkpoints had lower error
but appreciably less color, illustrating the need to inspect colorfulness.
These results support the added loss in this experiment; they do not prove
that the existing artifacts have one particular architectural cause.
### Broader data
A third GPU arm added 5,382 eligible COCO training photos to the 11,943
Imagenette photos. The first two pinned COCO training shards contain 5,864
photos; 200 were reserved for validation and the remaining photos were
filtered using the same low-chroma threshold. The combined training pool has
17,325 photos. The run used the classification objective and the same 1,122
updates and batch size as the GPU control, approximately 2.07 epochs over
the larger pool. The original color vocabulary and class weights stayed fixed.
The COCO training-validation set is distinct from the 100 COCO-val photos
used as an external probe. Image-ID intersection between the training shards
and that external sample is zero. The COCO revision is
`26ddc382fe75dfc2a0655b5977e296ea10efebce`. The actual indices, hashes and
training manifests are retained. Broader-data training was exploratory;
this is not a claim that a small COCO subset is sufficient for production.
## What the diagnostics establish
The constant-luminance and shifted-image probes show spatial sensitivity.
Across ten validation images, shifting by one pixel changed predicted chroma
by 2.31 Lab units on average in the central region; an eight-pixel shift
changed it by 1.45. Flat-input responses also contain spatial color variation.
This does not isolate a causal layer: padding, pooling, dilations, upsampling
and learned features can all contribute.
[Odena et al.](https://distill.pub/2016/deconv-checkerboard/) motivates testing
upsampling effects, but the current kernel-2/stride-2 layers do not have the
classic uneven-overlap configuration. Replacing them without a controlled
training comparison is not established here as a fix.
The current context path already covers hundreds of input pixels; simply
increasing dilation is not a demonstrated remedy. The remaining wrong hues
also point to a semantic problem that spatial filtering cannot solve. The
[original colorization research](https://richzhang.github.io/colorization/)
discusses dataset bias despite much larger training coverage, and
[DDColor](https://arxiv.org/abs/2212.11613) motivates semantic and multiscale
features. Semantic teacher-to-student training or a pretrained compact
encoder remain untested architectural alternatives here, not completed work.
## Evaluation limits
The 200-image Imagenette test is disjoint from the previous 100-image test
and the 50-image validation set. It comes from the historical seed-0 5%
holdout. Exposure during older upstream training remains unknown. The 100
COCO-val images are the first viewer rows, a convenience sample with 72
object categories represented, not a population-representative random test.
Chroma error measures closeness to one reference colorization; plausible
alternative colors may score poorly. Excess-edge measures are heuristics,
not counts of visible blotches or human quality ratings. Grayscale-source
references are separately counted and color-source-only summaries retained.
Qualitative examples include known failures, not only favorable results.
The additional 200-photo COCO training-validation comparison uses the reserved
indices, with 256px JPEG95 copies for transfer. Every candidate receives the
same decoded copies. These images are used for selecting weights and are
not reported as an independent final test.
## Execution and reproducibility
The custom plugin named Huggingface still exposed no tools. The original
Hugging Face connection supplied Jobs and read access but lacked write-repos.
A 17 MiB checksum-verified log roundtrip established a fallback artifact
transport. Larger two-checkpoint archives exceeded the tool's 64 MiB response
frame because the response repeats log content; CPU helpers exported one
checkpoint at a time. Explicit log-tail limits avoided server truncation.
All recovered archives are checked by size, SHA-256 and ZIP CRC.
Two initial GPU starts stopped on nonfinite FP16 gradients before completing
training. BF16 restarts completed with finite-gradient guards enabled. Those
failed starts are recorded separately from the five completed training runs.
The three GPU jobs were `6aaf8ceb52d0dbd7f1d73094`,
`6aaf8cf752d0dbd7f1d7309a`, and `6aaf8dad51992417dfccc557`.
Both best and last learned checkpoints are retained. GPU optimizer state was
not exported; those checkpoints support a new-optimizer fine-tune, not exact
training resume. CPU optimizer states are retained. Source, configurations,
environments, priors, split definitions and per-image results accompany them.
The corrected decoder and vocabulary guards continue to prevent the earlier
silent mapping mismatch.
|