File size: 13,661 Bytes
704fa80
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
# Mini U-Net Colorizer: spatial artifacts and further training

20 September 2026. This study starts from the **bin-mapping repair**, not the
broken Hub checkpoint. Improvements below must not be conflated with the
previous repair's 55.5% reduction in chroma error.

## Selected release and final results

**Selected:** mixed-data GPU checkpoint at update 748, from the completed
1,122-update run. Selection used 50 Imagenette plus 200 reserved COCO training
images, with equal weight for each domain's relative chroma error. A candidate
had to retain at least 90% of baseline predicted chroma in both domains and
avoid worse fine excess-edge scores. CPU and GPU spatial-loss candidates
failed the COCO color-retention gate (87.0% and 89.0%). The selected checkpoint
retained 96.5% and 97.9%, respectively, with a 2.66% improvement in the
selection score. No final-test results were used to change weights or decoding.

All comparisons below start from the already repaired baseline. Lower error
and excess-edge scores are better. Chroma ratio is predicted/reference mean
chroma; it is not a realism score.

| Test sample | Pipeline | Chroma error | Chroma ratio | Fine excess edge | Coarse excess edge |
|---|---|---:|---:|---:|---:|
| Imagenette 200 | Previous repair, raw | 13.318 | 0.965 | 0.5395 | 1.1962 |
| Imagenette 200 | Previous repair, guided8 | 13.077 | 0.943 | 0.0829 | 0.5948 |
| Imagenette 200 | New weights, raw | 13.056 | 0.919 | 0.5188 | 1.0782 |
| Imagenette 200 | **New weights, guided8** | **12.847** | **0.899** | **0.0765** | **0.5354** |
| COCO-val 100 | Previous repair, raw | 15.266 | 0.981 | 0.5896 | 1.4608 |
| COCO-val 100 | Previous repair, guided8 | 14.992 | 0.959 | 0.0990 | 0.7479 |
| COCO-val 100 | New weights, raw | 14.771 | 0.985 | 0.5813 | 1.3267 |
| COCO-val 100 | **New weights, guided8** | **14.532** | **0.966** | **0.0923** | **0.6727** |

Relative to the previous raw repair, the complete pipeline improves chroma
error by **3.54% and 4.81%**, fine excess edges by **85.83% and 84.35%**, and
coarse excess edges by **55.24% and 53.95%**, on Imagenette and COCO respectively.
These are measured proxies, not percentages of visible blotches eliminated.
Mean error improves on 142/200 and 76/100 images. Paired-image bootstrap 95%
intervals for mean error reduction are [0.308, 0.639] and [0.362, 1.089] Lab units.

There is a learned improvement beyond postprocessing: raw weights improve
error by 1.97% and 3.25%; with the same guided8 decoder on both old and new
weights, improvement is 1.76% and 3.07%. The corresponding guided-to-guided
paired intervals are [0.060, 0.398] and [0.091, 0.812]. Per-image results,
10,000-resample bootstrap calculations and low-chroma-source summaries are
provided in `reports/round2/final_results.json` and the final evaluation JSONs.

Qualitative checks retain failures. The coffee image has substantially fewer
fine rainbow patches but still contains wrong broad hues. The cat remains
undercolored. The rocket's color error worsens from 21.94 to 22.86 despite
smoothing. Four fixed probes and one fixed example per Imagenette class are
shown in `final_visual.png` and `final_test_comparison.png`; they were not
selected for favorable outcomes. The held-out quantitative pipeline resizes
images to 256x256; the four application probes exercise the aspect-preserving,
full-resolution output wrapper. Their metrics are not directly pooled.

The result is a useful **app-testing candidate**, not evidence that all
blotches are removed or that arbitrary real-world photos are production-ready.

## Deployment verification

The selected safetensors file has SHA-256
`0e4c417375684a044860f8af3ac3a2fb44e1a5729254ca33abda758ea71aea6e`.
All 65 learned parameter tensors changed from the repaired baseline; the
verified color-bin buffer is identical. Parameter count remains 3,968,892.

An ONNX opset-17 graph includes the learned model, temperature-0.38 decoding
and guided8 filtering. It is 15,885,757 bytes and accepts dynamic spatial sizes
and batches. Four numerical shape checks passed, with maximum Lab-ab difference
below 0.000009 against PyTorch. Three image-wrapper comparisons preserved image
dimensions and differed by at most one 8-bit RGB level. Ten code tests passed.
The wrapper and graph input/output contract are documented in `DEPLOYMENT.md`.

On this shared CPU host, ONNX Runtime 1.30.0 with two threads measured median
909ms and p90 1239ms for a 256x256 input (20 iterations after five warmups).
Peak process RSS was 508MiB. This includes network, decoding and filtering,
excludes file loading/full-resolution color conversion, and is not a browser
or mobile latency guarantee. The graph is not quantized.

Repository publication remains blocked: only the original Hugging Face
connection is callable, with Jobs/read scopes and no write-repos scope.
Previous direct and PR upload attempts returned 403; rechecking scopes at
completion confirmed the same restriction. Neither main nor stable changed.
The complete release, source, evidence and atomic main-upload helper are
preserved for manual publication. No additional training is left running.

## Questions and evidence

Three interventions were tested: spatial decoding, supervised spatial loss,
and broader training data. Five training runs completed, including three L4
GPU runs. Every model retains 3,968,892 learned parameters and the same 236
color-bin meanings. Model selection uses validation images, with separate
Imagenette and COCO external checks. Final quantitative results and the selected release follow below.

### Spatial decoding

Thirteen alternatives were compared on 50 validation images: the repaired
baseline; temperatures 0.6 and 0.8; logit pooling by factors 2, 4 and 8;
luminance-guided filtering with radii 4, 8 and 16; horizontal-flip averaging;
flip averaging plus radius-4 filtering; pooling plus filtering; and a gray
control. The gray control achieves zero discontinuity scores, demonstrating
why these scores cannot select a model by themselves.

The [guided-filter paper](https://people.csail.mit.edu/kaiming/eccv10/index.html)
motivates a local linear filter whose coefficients depend on luminance. This
implementation filters predicted Lab chroma using only input luminance;
reference colors are never passed into inference. It can smooth spurious
color changes while retaining luminance boundaries. Same-luminance color
boundaries remain a limitation.

Radius 8 was nominated as the balanced single-pass mode using validation
results and images. Radius 16 visibly removes more genuine small color
details, despite better artifact scores. Flip averaging plus radius 4 is an
optional two-pass mode. Temperature stays 0.38 and saturation stays 1.0.

On the additional 200-image Imagenette test, radius-8 filtering alone changed
mean chroma error from 13.318 to 13.077, fine excess color edge from 0.5395 to
0.0829, and coarse excess edge from 1.1962 to 0.5948. Mean predicted chroma
relative to reference changed from 0.965 to 0.943. Error improved on 179/200
images; paired image bootstrap 95% interval for the mean improvement was
[0.205, 0.276] Lab units.

On 100 COCO-val photos, filtering alone changed error from 15.266 to 14.992,
fine excess edge from 0.5896 to 0.0990, and coarse excess edge from 1.4608 to
0.7479. Chroma ratio changed from 0.981 to 0.959. Error improved on 95/100
images; the corresponding interval was [0.233, 0.314]. These intervals describe
the image samples, not performance across arbitrary deployment photos.

### Learned spatial consistency

The added loss matches chroma gradients to reference gradients at scales
1, 4 and 16. Chroma is normalized by 110; the gradient differences use smooth
L1 with beta 0.05, averaged over scales and horizontal/vertical directions.
Its coefficient is 10 alongside the original weighted classification loss.
This is supervised transition matching, rather than a requirement that every
region become uniformly gray.

The CPU pair used 400 updates, batch 2, learning rate 2e-6 and seed 0. The
GPU pair used 1,122 updates, batch 32, learning rate 1e-5 and seed 123: three
complete passes through 11,943 training photos. Within each pair, the only
training-objective difference is the spatial loss. Both use frozen BatchNorm
statistics, identical data order, fixed audited color weights, and cosine
learning-rate decay. CPU-to-GPU differences also change batch, precision,
seed and training duration and are not a single-factor comparison.

At the final CPU checkpoint, classification-only versus spatial loss produced
validation error 12.384 versus 12.149, CE 2.8076 versus 2.7918, and chroma
ratio 0.962 versus 0.923. At the final GPU checkpoint, error was 12.735 versus
12.326 and CE 2.8293 versus 2.8091. Earlier CPU checkpoints had lower error
but appreciably less color, illustrating the need to inspect colorfulness.

These results support the added loss in this experiment; they do not prove
that the existing artifacts have one particular architectural cause.

### Broader data

A third GPU arm added 5,382 eligible COCO training photos to the 11,943
Imagenette photos. The first two pinned COCO training shards contain 5,864
photos; 200 were reserved for validation and the remaining photos were
filtered using the same low-chroma threshold. The combined training pool has
17,325 photos. The run used the classification objective and the same 1,122
updates and batch size as the GPU control, approximately 2.07 epochs over
the larger pool. The original color vocabulary and class weights stayed fixed.

The COCO training-validation set is distinct from the 100 COCO-val photos
used as an external probe. Image-ID intersection between the training shards
and that external sample is zero. The COCO revision is
`26ddc382fe75dfc2a0655b5977e296ea10efebce`. The actual indices, hashes and
training manifests are retained. Broader-data training was exploratory;
this is not a claim that a small COCO subset is sufficient for production.

## What the diagnostics establish

The constant-luminance and shifted-image probes show spatial sensitivity.
Across ten validation images, shifting by one pixel changed predicted chroma
by 2.31 Lab units on average in the central region; an eight-pixel shift
changed it by 1.45. Flat-input responses also contain spatial color variation.
This does not isolate a causal layer: padding, pooling, dilations, upsampling
and learned features can all contribute.

[Odena et al.](https://distill.pub/2016/deconv-checkerboard/) motivates testing
upsampling effects, but the current kernel-2/stride-2 layers do not have the
classic uneven-overlap configuration. Replacing them without a controlled
training comparison is not established here as a fix.

The current context path already covers hundreds of input pixels; simply
increasing dilation is not a demonstrated remedy. The remaining wrong hues
also point to a semantic problem that spatial filtering cannot solve. The
[original colorization research](https://richzhang.github.io/colorization/)
discusses dataset bias despite much larger training coverage, and
[DDColor](https://arxiv.org/abs/2212.11613) motivates semantic and multiscale
features. Semantic teacher-to-student training or a pretrained compact
encoder remain untested architectural alternatives here, not completed work.

## Evaluation limits

The 200-image Imagenette test is disjoint from the previous 100-image test
and the 50-image validation set. It comes from the historical seed-0 5%
holdout. Exposure during older upstream training remains unknown. The 100
COCO-val images are the first viewer rows, a convenience sample with 72
object categories represented, not a population-representative random test.

Chroma error measures closeness to one reference colorization; plausible
alternative colors may score poorly. Excess-edge measures are heuristics,
not counts of visible blotches or human quality ratings. Grayscale-source
references are separately counted and color-source-only summaries retained.
Qualitative examples include known failures, not only favorable results.

The additional 200-photo COCO training-validation comparison uses the reserved
indices, with 256px JPEG95 copies for transfer. Every candidate receives the
same decoded copies. These images are used for selecting weights and are
not reported as an independent final test.

## Execution and reproducibility

The custom plugin named Huggingface still exposed no tools. The original
Hugging Face connection supplied Jobs and read access but lacked write-repos.
A 17 MiB checksum-verified log roundtrip established a fallback artifact
transport. Larger two-checkpoint archives exceeded the tool's 64 MiB response
frame because the response repeats log content; CPU helpers exported one
checkpoint at a time. Explicit log-tail limits avoided server truncation.
All recovered archives are checked by size, SHA-256 and ZIP CRC.

Two initial GPU starts stopped on nonfinite FP16 gradients before completing
training. BF16 restarts completed with finite-gradient guards enabled. Those
failed starts are recorded separately from the five completed training runs.
The three GPU jobs were `6aaf8ceb52d0dbd7f1d73094`,
`6aaf8cf752d0dbd7f1d7309a`, and `6aaf8dad51992417dfccc557`.

Both best and last learned checkpoints are retained. GPU optimizer state was
not exported; those checkpoints support a new-optimizer fine-tune, not exact
training resume. CPU optimizer states are retained. Source, configurations,
environments, priors, split definitions and per-image results accompany them.
The corrected decoder and vocabulary guards continue to prevent the earlier
silent mapping mismatch.