Document round6 selected model, fresh evaluation and limitations
Browse files
experiments/round6-20260928/RESULTS.md
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Round 6: pretrained semantic palette colorizer
|
| 2 |
+
|
| 3 |
+
Status: **promising experimental release candidate; not approved as a completed production model**. The selected checkpoint is `eligible_candidate/` (palette decoder, step 9000, 3,994,676 trainable parameters including the MobileNetV3 encoder). The older U-Net at the model repository root remains the current app checkpoint.
|
| 4 |
+
|
| 5 |
+
## Result
|
| 6 |
+
|
| 7 |
+
The previous attempt lowered local color differences by turning many areas gray. Round 6 instead used ImageNet-pretrained MobileNetV3 visual features and sixteen image-wide palette queries. It learned a single coherent DDColor artistic target per image, with pixel, color-distribution and target-boundary losses. A matched dense decoder trained with the same data and objective. The palette checkpoint at step 9000 passed the prespecified validation gates for both ordinary and film-style grayscale.
|
| 8 |
+
|
| 9 |
+
| Fresh COCO validation, ordinary grayscale | Root U-Net v2 | Palette step 9000 |
|
| 10 |
+
| --- | ---: | ---: |
|
| 11 |
+
| Lab ab error, lower | 14.888 | 13.095 |
|
| 12 |
+
| Blotch proxy, lower | 1.714 | 0.747 |
|
| 13 |
+
| Missed color on colorful targets, lower | 22.30% | 15.86% |
|
| 14 |
+
| Color coverage | 45.22% | 50.05% |
|
| 15 |
+
| Neutral spill on neutral regions, lower | 19.08% | 19.90% |
|
| 16 |
+
|
| 17 |
+
These results cover 272 color-bearing images of a 300-image COCO validation sample that was never used for this student’s training or checkpoint selection. Paired image bootstrap intervals for the new-minus-v2 differences exclude zero for error, blotches, missed color and coverage; neutral spill’s interval includes zero. Film-style input improved error, blotches, missed color and neutral spill, with little coverage change. See [paired comparison](posthoc_9000/paired_fresh_comparison.json), [full evaluation](posthoc_9000/REPORT.json), and [source/protocol](PROTOCOL.md).
|
| 18 |
+
|
| 19 |
+
[Fresh-image grid](palette_fresh_validation_grid.jpg) shows the step-6000 review candidate; its test numbers must not be attributed to step 9000. The [step-9000 probes](posthoc_9000/palette_9000.jpg) and the [archival portrait](posthoc_9000/migrant_mother.jpg) and [industrial photograph](posthoc_9000/power_house_mechanic.jpg) show the selected checkpoint. On those archival images, visually plausible colors cannot be verified as historical truth.
|
| 20 |
+
|
| 21 |
+
## Remaining problems
|
| 22 |
+
|
| 23 |
+
The moon probe is wrongly pink/red; other ambiguous scenes can receive a coherent but incorrect hue. Two archival examples and the 20-image grid are insufficient for production certification. The pretrained encoder and teacher have possible upstream evaluation overlap; exact training duplicates were excluded from the fresh subset, while near duplicates were not screened. The blotch score measures excess color differences in target-flat regions and does not replace human review.
|
| 24 |
+
|
| 25 |
+
## Deployment
|
| 26 |
+
|
| 27 |
+
The candidate is saved and checksum-verified on `main`; the `stable` branch and root checkpoint are unchanged. Its loader in `source/semantic_model.py` needs TorchVision, and the current Space’s U-Net loader cannot use this checkpoint. Integrate and test the new architecture in the Space before changing the root release. A later round should specifically address semantically neutral objects such as lunar terrain while preserving color on genuine black-and-white portraits. The parameter limit remains below four million.
|
experiments/round6-20260928/SHA256SUMS.json
CHANGED
|
@@ -1,5 +1,6 @@
|
|
| 1 |
{
|
| 2 |
"PROTOCOL.md": "9b8094cf5cc2c785409b4a9850da742e5fab193e7ab96bc4a5fbd524f4327aa8",
|
|
|
|
| 3 |
"baseline_validation.json": "087e92c4f781275a8687cf6f9484ea02588add5ff1924d32d4a01b41a7575b02",
|
| 4 |
"dense/initial/README.md": "67b6b8e7697e8a64f150dc87760d54081af5b0f20d826b6d685c8fc9c7a121ef",
|
| 5 |
"dense/initial/config.json": "b45b92885f4715a02034212261db30ba48d6ed39af57c3a1d08a3b5e464f1fdc",
|
|
@@ -68,7 +69,7 @@
|
|
| 68 |
"dense_review_validation_grid.jpg": "b06edf59b6dda0f6d11b4ef2b4d5aef9271cb02471136c630535491b26c74716",
|
| 69 |
"dense_test.json": "b8f8ce4e611fee81ac1015cd9ad0cebeb5b4208d29180e3c4c5ba7be26e1bc2b",
|
| 70 |
"dense_test_validation_grid.jpg": "a05d8dca818bd93b3fce82fd761c8ef2e025bcf8c12794bc65b71e4ad8b9d499",
|
| 71 |
-
"eligible_candidate/README.md": "
|
| 72 |
"eligible_candidate/config.json": "eaeb49b39f9116bfed0d954e81852dc93476298752787dba88a63c0eec97ccec",
|
| 73 |
"eligible_candidate/model.safetensors": "ec1f27d74533adc83f7ab3639a091fc4d8738a434dafc7d172c7873c28a9e715",
|
| 74 |
"fresh_holdout_manifest.json": "3cb81e88dcd6c2b2c7c5a1fa69de6eeb86c5da9fe6e659918311cc80fcce64f1",
|
|
|
|
| 1 |
{
|
| 2 |
"PROTOCOL.md": "9b8094cf5cc2c785409b4a9850da742e5fab193e7ab96bc4a5fbd524f4327aa8",
|
| 3 |
+
"RESULTS.md": "1286c3f1a1ecffac89501f5dff3625608cfd71a6f4e596a5368a9c38244b9e6d",
|
| 4 |
"baseline_validation.json": "087e92c4f781275a8687cf6f9484ea02588add5ff1924d32d4a01b41a7575b02",
|
| 5 |
"dense/initial/README.md": "67b6b8e7697e8a64f150dc87760d54081af5b0f20d826b6d685c8fc9c7a121ef",
|
| 6 |
"dense/initial/config.json": "b45b92885f4715a02034212261db30ba48d6ed39af57c3a1d08a3b5e464f1fdc",
|
|
|
|
| 69 |
"dense_review_validation_grid.jpg": "b06edf59b6dda0f6d11b4ef2b4d5aef9271cb02471136c630535491b26c74716",
|
| 70 |
"dense_test.json": "b8f8ce4e611fee81ac1015cd9ad0cebeb5b4208d29180e3c4c5ba7be26e1bc2b",
|
| 71 |
"dense_test_validation_grid.jpg": "a05d8dca818bd93b3fce82fd761c8ef2e025bcf8c12794bc65b71e4ad8b9d499",
|
| 72 |
+
"eligible_candidate/README.md": "db08deccf6ed0920531d5634698e99dab74a2644629235a8d2fe3e18d78382d2",
|
| 73 |
"eligible_candidate/config.json": "eaeb49b39f9116bfed0d954e81852dc93476298752787dba88a63c0eec97ccec",
|
| 74 |
"eligible_candidate/model.safetensors": "ec1f27d74533adc83f7ab3639a091fc4d8738a434dafc7d172c7873c28a9e715",
|
| 75 |
"fresh_holdout_manifest.json": "3cb81e88dcd6c2b2c7c5a1fa69de6eeb86c5da9fe6e659918311cc80fcce64f1",
|
experiments/round6-20260928/eligible_candidate/README.md
CHANGED
|
@@ -1,5 +1,19 @@
|
|
| 1 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
|
| 3 |
-
|
| 4 |
|
| 5 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
pipeline_tag: image-to-image
|
| 4 |
+
tags:
|
| 5 |
+
- colorization
|
| 6 |
+
- image-to-image
|
| 7 |
+
- experimental
|
| 8 |
+
- mobilenetv3
|
| 9 |
+
---
|
| 10 |
|
| 11 |
+
# Mini colorizer: semantic palette candidate (round 6)
|
| 12 |
|
| 13 |
+
This is an **experimental checkpoint** at step 9000, selected on development validation. It has **3,994,676 learned parameters including the ImageNet-pretrained MobileNetV3 encoder**. The classifier is removed. It predicts Lab `ab` colors from normalized Lab lightness, combining sixteen image-wide palette colors with semantic assignment masks and a small bounded local residual. Color is plausible, not recovered historical truth.
|
| 14 |
+
|
| 15 |
+
The training set was 16,230 filtered images from pinned Imagenette and COCO training Parquet. Target color came from pinned `piddnad/ddcolor_artistic` output, one coherent prediction per image. The encoder and teacher may have seen overlapping images upstream. Training source and exact revisions are in `../PROTOCOL.md` and `../provenance.json`. This model is not the DDColor teacher and does not include its weights at inference.
|
| 16 |
+
|
| 17 |
+
On 272 colorful images in a fresh 300-image COCO validation sample, the model improved Lab ab error (14.89 to 13.10), excess flat-region color changes (1.71 to 0.75), and missed color (22.3% to 15.9%) relative to the root U-Net. Color coverage increased from 45.2% to 50.0%. Neutral-region spill was 19.1% to 19.9%, without a conclusive paired difference. These are proxy metrics. See `../RESULTS.md` and `../posthoc_9000/REPORT.json`.
|
| 18 |
+
|
| 19 |
+
**Known failure:** Lunar terrain is assigned pink/red in the visual probe. Other inherently neutral or ambiguous objects can get incorrect hues. Actual grayscale historical photos have no color ground truth; two archival examples are visual checks only. The current Hugging Face Space expects the previous U-Net architecture, so use `../source/semantic_model.py` and `load_semantic` for this checkpoint. It is not yet the app’s root release.
|