|
Download RESEARCH_ROUND2.md from User-2468/mini-unet-colorizer: direct link, hf CLI and curl.
- Browser
- Download file 13.7 kB
-
https://huggingface.co/User-2468/mini-unet-colorizer/resolve/955113ecfd7f893e3aa1e591e27824fe603d3316/RESEARCH_ROUND2.md
- Command line
-
hf download hf://User-2468/mini-unet-colorizer@955113ecfd7f893e3aa1e591e27824fe603d3316/RESEARCH_ROUND2.md
-
curl -L -o RESEARCH_ROUND2.md https://huggingface.co/User-2468/mini-unet-colorizer/resolve/955113ecfd7f893e3aa1e591e27824fe603d3316/RESEARCH_ROUND2.md
13.7 kB
| # Mini U-Net Colorizer: spatial artifacts and further training | |
| 20 September 2026. This study starts from the **bin-mapping repair**, not the | |
| broken Hub checkpoint. Improvements below must not be conflated with the | |
| previous repair's 55.5% reduction in chroma error. | |
| ## Selected release and final results | |
| **Selected:** mixed-data GPU checkpoint at update 748, from the completed | |
| 1,122-update run. Selection used 50 Imagenette plus 200 reserved COCO training | |
| images, with equal weight for each domain's relative chroma error. A candidate | |
| had to retain at least 90% of baseline predicted chroma in both domains and | |
| avoid worse fine excess-edge scores. CPU and GPU spatial-loss candidates | |
| failed the COCO color-retention gate (87.0% and 89.0%). The selected checkpoint | |
| retained 96.5% and 97.9%, respectively, with a 2.66% improvement in the | |
| selection score. No final-test results were used to change weights or decoding. | |
| All comparisons below start from the already repaired baseline. Lower error | |
| and excess-edge scores are better. Chroma ratio is predicted/reference mean | |
| chroma; it is not a realism score. | |
| | Test sample | Pipeline | Chroma error | Chroma ratio | Fine excess edge | Coarse excess edge | | |
| |---|---|---:|---:|---:|---:| | |
| | Imagenette 200 | Previous repair, raw | 13.318 | 0.965 | 0.5395 | 1.1962 | | |
| | Imagenette 200 | Previous repair, guided8 | 13.077 | 0.943 | 0.0829 | 0.5948 | | |
| | Imagenette 200 | New weights, raw | 13.056 | 0.919 | 0.5188 | 1.0782 | | |
| | Imagenette 200 | **New weights, guided8** | **12.847** | **0.899** | **0.0765** | **0.5354** | | |
| | COCO-val 100 | Previous repair, raw | 15.266 | 0.981 | 0.5896 | 1.4608 | | |
| | COCO-val 100 | Previous repair, guided8 | 14.992 | 0.959 | 0.0990 | 0.7479 | | |
| | COCO-val 100 | New weights, raw | 14.771 | 0.985 | 0.5813 | 1.3267 | | |
| | COCO-val 100 | **New weights, guided8** | **14.532** | **0.966** | **0.0923** | **0.6727** | | |
| Relative to the previous raw repair, the complete pipeline improves chroma | |
| error by **3.54% and 4.81%**, fine excess edges by **85.83% and 84.35%**, and | |
| coarse excess edges by **55.24% and 53.95%**, on Imagenette and COCO respectively. | |
| These are measured proxies, not percentages of visible blotches eliminated. | |
| Mean error improves on 142/200 and 76/100 images. Paired-image bootstrap 95% | |
| intervals for mean error reduction are [0.308, 0.639] and [0.362, 1.089] Lab units. | |
| There is a learned improvement beyond postprocessing: raw weights improve | |
| error by 1.97% and 3.25%; with the same guided8 decoder on both old and new | |
| weights, improvement is 1.76% and 3.07%. The corresponding guided-to-guided | |
| paired intervals are [0.060, 0.398] and [0.091, 0.812]. Per-image results, | |
| 10,000-resample bootstrap calculations and low-chroma-source summaries are | |
| provided in `reports/round2/final_results.json` and the final evaluation JSONs. | |
| Qualitative checks retain failures. The coffee image has substantially fewer | |
| fine rainbow patches but still contains wrong broad hues. The cat remains | |
| undercolored. The rocket's color error worsens from 21.94 to 22.86 despite | |
| smoothing. Four fixed probes and one fixed example per Imagenette class are | |
| shown in `final_visual.png` and `final_test_comparison.png`; they were not | |
| selected for favorable outcomes. The held-out quantitative pipeline resizes | |
| images to 256x256; the four application probes exercise the aspect-preserving, | |
| full-resolution output wrapper. Their metrics are not directly pooled. | |
| The result is a useful **app-testing candidate**, not evidence that all | |
| blotches are removed or that arbitrary real-world photos are production-ready. | |
| ## Deployment verification | |
| The selected safetensors file has SHA-256 | |
| `0e4c417375684a044860f8af3ac3a2fb44e1a5729254ca33abda758ea71aea6e`. | |
| All 65 learned parameter tensors changed from the repaired baseline; the | |
| verified color-bin buffer is identical. Parameter count remains 3,968,892. | |
| An ONNX opset-17 graph includes the learned model, temperature-0.38 decoding | |
| and guided8 filtering. It is 15,885,757 bytes and accepts dynamic spatial sizes | |
| and batches. Four numerical shape checks passed, with maximum Lab-ab difference | |
| below 0.000009 against PyTorch. Three image-wrapper comparisons preserved image | |
| dimensions and differed by at most one 8-bit RGB level. Ten code tests passed. | |
| The wrapper and graph input/output contract are documented in `DEPLOYMENT.md`. | |
| On this shared CPU host, ONNX Runtime 1.30.0 with two threads measured median | |
| 909ms and p90 1239ms for a 256x256 input (20 iterations after five warmups). | |
| Peak process RSS was 508MiB. This includes network, decoding and filtering, | |
| excludes file loading/full-resolution color conversion, and is not a browser | |
| or mobile latency guarantee. The graph is not quantized. | |
| Repository publication remains blocked: only the original Hugging Face | |
| connection is callable, with Jobs/read scopes and no write-repos scope. | |
| Previous direct and PR upload attempts returned 403; rechecking scopes at | |
| completion confirmed the same restriction. Neither main nor stable changed. | |
| The complete release, source, evidence and atomic main-upload helper are | |
| preserved for manual publication. No additional training is left running. | |
| ## Questions and evidence | |
| Three interventions were tested: spatial decoding, supervised spatial loss, | |
| and broader training data. Five training runs completed, including three L4 | |
| GPU runs. Every model retains 3,968,892 learned parameters and the same 236 | |
| color-bin meanings. Model selection uses validation images, with separate | |
| Imagenette and COCO external checks. Final quantitative results and the selected release follow below. | |
| ### Spatial decoding | |
| Thirteen alternatives were compared on 50 validation images: the repaired | |
| baseline; temperatures 0.6 and 0.8; logit pooling by factors 2, 4 and 8; | |
| luminance-guided filtering with radii 4, 8 and 16; horizontal-flip averaging; | |
| flip averaging plus radius-4 filtering; pooling plus filtering; and a gray | |
| control. The gray control achieves zero discontinuity scores, demonstrating | |
| why these scores cannot select a model by themselves. | |
| The [guided-filter paper](https://people.csail.mit.edu/kaiming/eccv10/index.html) | |
| motivates a local linear filter whose coefficients depend on luminance. This | |
| implementation filters predicted Lab chroma using only input luminance; | |
| reference colors are never passed into inference. It can smooth spurious | |
| color changes while retaining luminance boundaries. Same-luminance color | |
| boundaries remain a limitation. | |
| Radius 8 was nominated as the balanced single-pass mode using validation | |
| results and images. Radius 16 visibly removes more genuine small color | |
| details, despite better artifact scores. Flip averaging plus radius 4 is an | |
| optional two-pass mode. Temperature stays 0.38 and saturation stays 1.0. | |
| On the additional 200-image Imagenette test, radius-8 filtering alone changed | |
| mean chroma error from 13.318 to 13.077, fine excess color edge from 0.5395 to | |
| 0.0829, and coarse excess edge from 1.1962 to 0.5948. Mean predicted chroma | |
| relative to reference changed from 0.965 to 0.943. Error improved on 179/200 | |
| images; paired image bootstrap 95% interval for the mean improvement was | |
| [0.205, 0.276] Lab units. | |
| On 100 COCO-val photos, filtering alone changed error from 15.266 to 14.992, | |
| fine excess edge from 0.5896 to 0.0990, and coarse excess edge from 1.4608 to | |
| 0.7479. Chroma ratio changed from 0.981 to 0.959. Error improved on 95/100 | |
| images; the corresponding interval was [0.233, 0.314]. These intervals describe | |
| the image samples, not performance across arbitrary deployment photos. | |
| ### Learned spatial consistency | |
| The added loss matches chroma gradients to reference gradients at scales | |
| 1, 4 and 16. Chroma is normalized by 110; the gradient differences use smooth | |
| L1 with beta 0.05, averaged over scales and horizontal/vertical directions. | |
| Its coefficient is 10 alongside the original weighted classification loss. | |
| This is supervised transition matching, rather than a requirement that every | |
| region become uniformly gray. | |
| The CPU pair used 400 updates, batch 2, learning rate 2e-6 and seed 0. The | |
| GPU pair used 1,122 updates, batch 32, learning rate 1e-5 and seed 123: three | |
| complete passes through 11,943 training photos. Within each pair, the only | |
| training-objective difference is the spatial loss. Both use frozen BatchNorm | |
| statistics, identical data order, fixed audited color weights, and cosine | |
| learning-rate decay. CPU-to-GPU differences also change batch, precision, | |
| seed and training duration and are not a single-factor comparison. | |
| At the final CPU checkpoint, classification-only versus spatial loss produced | |
| validation error 12.384 versus 12.149, CE 2.8076 versus 2.7918, and chroma | |
| ratio 0.962 versus 0.923. At the final GPU checkpoint, error was 12.735 versus | |
| 12.326 and CE 2.8293 versus 2.8091. Earlier CPU checkpoints had lower error | |
| but appreciably less color, illustrating the need to inspect colorfulness. | |
| These results support the added loss in this experiment; they do not prove | |
| that the existing artifacts have one particular architectural cause. | |
| ### Broader data | |
| A third GPU arm added 5,382 eligible COCO training photos to the 11,943 | |
| Imagenette photos. The first two pinned COCO training shards contain 5,864 | |
| photos; 200 were reserved for validation and the remaining photos were | |
| filtered using the same low-chroma threshold. The combined training pool has | |
| 17,325 photos. The run used the classification objective and the same 1,122 | |
| updates and batch size as the GPU control, approximately 2.07 epochs over | |
| the larger pool. The original color vocabulary and class weights stayed fixed. | |
| The COCO training-validation set is distinct from the 100 COCO-val photos | |
| used as an external probe. Image-ID intersection between the training shards | |
| and that external sample is zero. The COCO revision is | |
| `26ddc382fe75dfc2a0655b5977e296ea10efebce`. The actual indices, hashes and | |
| training manifests are retained. Broader-data training was exploratory; | |
| this is not a claim that a small COCO subset is sufficient for production. | |
| ## What the diagnostics establish | |
| The constant-luminance and shifted-image probes show spatial sensitivity. | |
| Across ten validation images, shifting by one pixel changed predicted chroma | |
| by 2.31 Lab units on average in the central region; an eight-pixel shift | |
| changed it by 1.45. Flat-input responses also contain spatial color variation. | |
| This does not isolate a causal layer: padding, pooling, dilations, upsampling | |
| and learned features can all contribute. | |
| [Odena et al.](https://distill.pub/2016/deconv-checkerboard/) motivates testing | |
| upsampling effects, but the current kernel-2/stride-2 layers do not have the | |
| classic uneven-overlap configuration. Replacing them without a controlled | |
| training comparison is not established here as a fix. | |
| The current context path already covers hundreds of input pixels; simply | |
| increasing dilation is not a demonstrated remedy. The remaining wrong hues | |
| also point to a semantic problem that spatial filtering cannot solve. The | |
| [original colorization research](https://richzhang.github.io/colorization/) | |
| discusses dataset bias despite much larger training coverage, and | |
| [DDColor](https://arxiv.org/abs/2212.11613) motivates semantic and multiscale | |
| features. Semantic teacher-to-student training or a pretrained compact | |
| encoder remain untested architectural alternatives here, not completed work. | |
| ## Evaluation limits | |
| The 200-image Imagenette test is disjoint from the previous 100-image test | |
| and the 50-image validation set. It comes from the historical seed-0 5% | |
| holdout. Exposure during older upstream training remains unknown. The 100 | |
| COCO-val images are the first viewer rows, a convenience sample with 72 | |
| object categories represented, not a population-representative random test. | |
| Chroma error measures closeness to one reference colorization; plausible | |
| alternative colors may score poorly. Excess-edge measures are heuristics, | |
| not counts of visible blotches or human quality ratings. Grayscale-source | |
| references are separately counted and color-source-only summaries retained. | |
| Qualitative examples include known failures, not only favorable results. | |
| The additional 200-photo COCO training-validation comparison uses the reserved | |
| indices, with 256px JPEG95 copies for transfer. Every candidate receives the | |
| same decoded copies. These images are used for selecting weights and are | |
| not reported as an independent final test. | |
| ## Execution and reproducibility | |
| The custom plugin named Huggingface still exposed no tools. The original | |
| Hugging Face connection supplied Jobs and read access but lacked write-repos. | |
| A 17 MiB checksum-verified log roundtrip established a fallback artifact | |
| transport. Larger two-checkpoint archives exceeded the tool's 64 MiB response | |
| frame because the response repeats log content; CPU helpers exported one | |
| checkpoint at a time. Explicit log-tail limits avoided server truncation. | |
| All recovered archives are checked by size, SHA-256 and ZIP CRC. | |
| Two initial GPU starts stopped on nonfinite FP16 gradients before completing | |
| training. BF16 restarts completed with finite-gradient guards enabled. Those | |
| failed starts are recorded separately from the five completed training runs. | |
| The three GPU jobs were `6aaf8ceb52d0dbd7f1d73094`, | |
| `6aaf8cf752d0dbd7f1d7309a`, and `6aaf8dad51992417dfccc557`. | |
| Both best and last learned checkpoints are retained. GPU optimizer state was | |
| not exported; those checkpoints support a new-optimizer fine-tune, not exact | |
| training resume. CPU optimizer states are retained. Source, configurations, | |
| environments, priors, split definitions and per-image results accompany them. | |
| The corrected decoder and vocabulary guards continue to prevent the earlier | |
| silent mapping mismatch. | |