User-2468 commited on
Commit
7b10485
·
verified ·
1 Parent(s): 6a6ed21

Document round6 selected model, fresh evaluation and limitations

Browse files
experiments/round6-20260928/RESULTS.md ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Round 6: pretrained semantic palette colorizer
2
+
3
+ Status: **promising experimental release candidate; not approved as a completed production model**. The selected checkpoint is `eligible_candidate/` (palette decoder, step 9000, 3,994,676 trainable parameters including the MobileNetV3 encoder). The older U-Net at the model repository root remains the current app checkpoint.
4
+
5
+ ## Result
6
+
7
+ The previous attempt lowered local color differences by turning many areas gray. Round 6 instead used ImageNet-pretrained MobileNetV3 visual features and sixteen image-wide palette queries. It learned a single coherent DDColor artistic target per image, with pixel, color-distribution and target-boundary losses. A matched dense decoder trained with the same data and objective. The palette checkpoint at step 9000 passed the prespecified validation gates for both ordinary and film-style grayscale.
8
+
9
+ | Fresh COCO validation, ordinary grayscale | Root U-Net v2 | Palette step 9000 |
10
+ | --- | ---: | ---: |
11
+ | Lab ab error, lower | 14.888 | 13.095 |
12
+ | Blotch proxy, lower | 1.714 | 0.747 |
13
+ | Missed color on colorful targets, lower | 22.30% | 15.86% |
14
+ | Color coverage | 45.22% | 50.05% |
15
+ | Neutral spill on neutral regions, lower | 19.08% | 19.90% |
16
+
17
+ These results cover 272 color-bearing images of a 300-image COCO validation sample that was never used for this student’s training or checkpoint selection. Paired image bootstrap intervals for the new-minus-v2 differences exclude zero for error, blotches, missed color and coverage; neutral spill’s interval includes zero. Film-style input improved error, blotches, missed color and neutral spill, with little coverage change. See [paired comparison](posthoc_9000/paired_fresh_comparison.json), [full evaluation](posthoc_9000/REPORT.json), and [source/protocol](PROTOCOL.md).
18
+
19
+ [Fresh-image grid](palette_fresh_validation_grid.jpg) shows the step-6000 review candidate; its test numbers must not be attributed to step 9000. The [step-9000 probes](posthoc_9000/palette_9000.jpg) and the [archival portrait](posthoc_9000/migrant_mother.jpg) and [industrial photograph](posthoc_9000/power_house_mechanic.jpg) show the selected checkpoint. On those archival images, visually plausible colors cannot be verified as historical truth.
20
+
21
+ ## Remaining problems
22
+
23
+ The moon probe is wrongly pink/red; other ambiguous scenes can receive a coherent but incorrect hue. Two archival examples and the 20-image grid are insufficient for production certification. The pretrained encoder and teacher have possible upstream evaluation overlap; exact training duplicates were excluded from the fresh subset, while near duplicates were not screened. The blotch score measures excess color differences in target-flat regions and does not replace human review.
24
+
25
+ ## Deployment
26
+
27
+ The candidate is saved and checksum-verified on `main`; the `stable` branch and root checkpoint are unchanged. Its loader in `source/semantic_model.py` needs TorchVision, and the current Space’s U-Net loader cannot use this checkpoint. Integrate and test the new architecture in the Space before changing the root release. A later round should specifically address semantically neutral objects such as lunar terrain while preserving color on genuine black-and-white portraits. The parameter limit remains below four million.
experiments/round6-20260928/SHA256SUMS.json CHANGED
@@ -1,5 +1,6 @@
1
  {
2
  "PROTOCOL.md": "9b8094cf5cc2c785409b4a9850da742e5fab193e7ab96bc4a5fbd524f4327aa8",
 
3
  "baseline_validation.json": "087e92c4f781275a8687cf6f9484ea02588add5ff1924d32d4a01b41a7575b02",
4
  "dense/initial/README.md": "67b6b8e7697e8a64f150dc87760d54081af5b0f20d826b6d685c8fc9c7a121ef",
5
  "dense/initial/config.json": "b45b92885f4715a02034212261db30ba48d6ed39af57c3a1d08a3b5e464f1fdc",
@@ -68,7 +69,7 @@
68
  "dense_review_validation_grid.jpg": "b06edf59b6dda0f6d11b4ef2b4d5aef9271cb02471136c630535491b26c74716",
69
  "dense_test.json": "b8f8ce4e611fee81ac1015cd9ad0cebeb5b4208d29180e3c4c5ba7be26e1bc2b",
70
  "dense_test_validation_grid.jpg": "a05d8dca818bd93b3fce82fd761c8ef2e025bcf8c12794bc65b71e4ad8b9d499",
71
- "eligible_candidate/README.md": "2508e51f14898f651c49adb184d1570b8adad3702313dee9bdac6fa0390d2bec",
72
  "eligible_candidate/config.json": "eaeb49b39f9116bfed0d954e81852dc93476298752787dba88a63c0eec97ccec",
73
  "eligible_candidate/model.safetensors": "ec1f27d74533adc83f7ab3639a091fc4d8738a434dafc7d172c7873c28a9e715",
74
  "fresh_holdout_manifest.json": "3cb81e88dcd6c2b2c7c5a1fa69de6eeb86c5da9fe6e659918311cc80fcce64f1",
 
1
  {
2
  "PROTOCOL.md": "9b8094cf5cc2c785409b4a9850da742e5fab193e7ab96bc4a5fbd524f4327aa8",
3
+ "RESULTS.md": "1286c3f1a1ecffac89501f5dff3625608cfd71a6f4e596a5368a9c38244b9e6d",
4
  "baseline_validation.json": "087e92c4f781275a8687cf6f9484ea02588add5ff1924d32d4a01b41a7575b02",
5
  "dense/initial/README.md": "67b6b8e7697e8a64f150dc87760d54081af5b0f20d826b6d685c8fc9c7a121ef",
6
  "dense/initial/config.json": "b45b92885f4715a02034212261db30ba48d6ed39af57c3a1d08a3b5e464f1fdc",
 
69
  "dense_review_validation_grid.jpg": "b06edf59b6dda0f6d11b4ef2b4d5aef9271cb02471136c630535491b26c74716",
70
  "dense_test.json": "b8f8ce4e611fee81ac1015cd9ad0cebeb5b4208d29180e3c4c5ba7be26e1bc2b",
71
  "dense_test_validation_grid.jpg": "a05d8dca818bd93b3fce82fd761c8ef2e025bcf8c12794bc65b71e4ad8b9d499",
72
+ "eligible_candidate/README.md": "db08deccf6ed0920531d5634698e99dab74a2644629235a8d2fe3e18d78382d2",
73
  "eligible_candidate/config.json": "eaeb49b39f9116bfed0d954e81852dc93476298752787dba88a63c0eec97ccec",
74
  "eligible_candidate/model.safetensors": "ec1f27d74533adc83f7ab3639a091fc4d8738a434dafc7d172c7873c28a9e715",
75
  "fresh_holdout_manifest.json": "3cb81e88dcd6c2b2c7c5a1fa69de6eeb86c5da9fe6e659918311cc80fcce64f1",
experiments/round6-20260928/eligible_candidate/README.md CHANGED
@@ -1,5 +1,19 @@
1
- # Experimental semantic colorizer
 
 
 
 
 
 
 
 
2
 
3
- Head: palette. Parameters: 3994676.
4
 
5
- Not approved for production. Requires semantic_model.py; incompatible with the old U-Net loader. Input is Lab lightness normalized to [-1,1]; output is Lab ab. See the run protocol, provenance, selection and visual comparisons. Predictions are plausible colors, not recovered historical truth.
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ pipeline_tag: image-to-image
4
+ tags:
5
+ - colorization
6
+ - image-to-image
7
+ - experimental
8
+ - mobilenetv3
9
+ ---
10
 
11
+ # Mini colorizer: semantic palette candidate (round 6)
12
 
13
+ This is an **experimental checkpoint** at step 9000, selected on development validation. It has **3,994,676 learned parameters including the ImageNet-pretrained MobileNetV3 encoder**. The classifier is removed. It predicts Lab `ab` colors from normalized Lab lightness, combining sixteen image-wide palette colors with semantic assignment masks and a small bounded local residual. Color is plausible, not recovered historical truth.
14
+
15
+ The training set was 16,230 filtered images from pinned Imagenette and COCO training Parquet. Target color came from pinned `piddnad/ddcolor_artistic` output, one coherent prediction per image. The encoder and teacher may have seen overlapping images upstream. Training source and exact revisions are in `../PROTOCOL.md` and `../provenance.json`. This model is not the DDColor teacher and does not include its weights at inference.
16
+
17
+ On 272 colorful images in a fresh 300-image COCO validation sample, the model improved Lab ab error (14.89 to 13.10), excess flat-region color changes (1.71 to 0.75), and missed color (22.3% to 15.9%) relative to the root U-Net. Color coverage increased from 45.2% to 50.0%. Neutral-region spill was 19.1% to 19.9%, without a conclusive paired difference. These are proxy metrics. See `../RESULTS.md` and `../posthoc_9000/REPORT.json`.
18
+
19
+ **Known failure:** Lunar terrain is assigned pink/red in the visual probe. Other inherently neutral or ambiguous objects can get incorrect hues. Actual grayscale historical photos have no color ground truth; two archival examples are visual checks only. The current Hugging Face Space expects the previous U-Net architecture, so use `../source/semantic_model.py` and `load_semantic` for this checkpoint. It is not yet the app’s root release.