User-2468 commited on
Commit
704fa80
·
verified ·
1 Parent(s): 6c47ea4

Update validated colorizer weights, spatial decoder and deployment artifacts

Browse files
.gitattributes CHANGED
@@ -37,3 +37,6 @@ sample.png filter=lfs diff=lfs merge=lfs -text
37
  eval_grid.png filter=lfs diff=lfs merge=lfs -text
38
  eval_grid_temp075.png filter=lfs diff=lfs merge=lfs -text
39
  eval_grid_temp038.png filter=lfs diff=lfs merge=lfs -text
 
 
 
 
37
  eval_grid.png filter=lfs diff=lfs merge=lfs -text
38
  eval_grid_temp075.png filter=lfs diff=lfs merge=lfs -text
39
  eval_grid_temp038.png filter=lfs diff=lfs merge=lfs -text
40
+ reports/round2/final_coco100.png filter=lfs diff=lfs merge=lfs -text
41
+ reports/round2/final_test_comparison.png filter=lfs diff=lfs merge=lfs -text
42
+ reports/round2/final_visual.png filter=lfs diff=lfs merge=lfs -text
DEPLOYMENT.md ADDED
@@ -0,0 +1,80 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Colorizer deployment
2
+
3
+ Use the release's inference wrapper or its ONNX graph to obtain the complete
4
+ improvement. The `.safetensors` file contains the learned colorizer; guided
5
+ decoding is implemented by `spatial.py` and is included in the ONNX graph.
6
+
7
+ ## Python / PyTorch
8
+
9
+ Run from the extracted release directory:
10
+
11
+ ```bash
12
+ python -m pip install -r requirements.txt
13
+ python inference.py --model . --output-dir colorized photo.jpg
14
+ ```
15
+
16
+ The default guided radius is 8 at the model's input resolution. Use
17
+ `--guided-radius 0` for raw predictions or `--guided-radius 16` for stronger
18
+ smoothing. Stronger smoothing can remove legitimate small color details.
19
+ `--flip-tta --guided-radius 4` is an optional two-pass mode. The released
20
+ ONNX graph is the single-pass radius-8 mode.
21
+
22
+ ```python
23
+ from PIL import Image
24
+ from model import load_model
25
+ from inference import colorize
26
+
27
+ model = load_model('.')
28
+ result = colorize(model, Image.open('photo.jpg'))
29
+ result.save('colorized.png')
30
+ ```
31
+
32
+ The wrapper handles EXIF orientation, preserves aspect ratio, bounds the
33
+ longest network-input side to 256 pixels, upsamples chroma to the original
34
+ oriented image dimensions and combines it with the original Lab luminance.
35
+ Final RGB conversion can clip colors outside the display gamut. Very large
36
+ inputs still require memory for full-resolution color conversion. There is
37
+ no video temporal-consistency guarantee.
38
+
39
+ ## ONNX without PyTorch
40
+
41
+ ```bash
42
+ python -m pip install -r requirements-onnx.txt
43
+ python colorize_onnx.py --model colorizer.onnx --output-dir colorized photo.jpg
44
+ ```
45
+
46
+ Input name: `luminance`, float32, shape `N x 1 x H x W`, values `L*/50 - 1`.
47
+ Output name: `chroma`, float32, shape `N x 2 x H x W`, Lab a and b values.
48
+ Height and width must each be at least 8. Dynamic shapes and batches are
49
+ supported. Prefer a longest input side of 256 to match the evaluated
50
+ operating point. Ordinary RGB values are not valid graph inputs.
51
+
52
+ The graph includes temperature-0.38 decoding and radius-8 guided filtering
53
+ with epsilon 0.001. It does not include file loading, EXIF handling, Lab
54
+ conversion, aspect-ratio resizing or final chroma upsampling; these are
55
+ implemented in `colorize_onnx.py`. The ONNX wrapper uses Pillow resizing,
56
+ whereas the PyTorch wrapper uses PyTorch interpolation. Their image-level
57
+ comparison is recorded in the release checks.
58
+
59
+ ## Publish to main
60
+
61
+ Authenticate normally on your own computer with repository write access:
62
+
63
+ ```bash
64
+ hf auth login
65
+ python upload_main.py --folder . --repo User-2468/mini-unet-colorizer
66
+ ```
67
+
68
+ This performs one upload to `main` with an optimistic-concurrency guard.
69
+ It also verifies that `stable` retains its pre-upload revision. The app can
70
+ keep using `stable` until you choose to switch it to the tested release.
71
+ The current chat connection has Jobs/read access but no repository write
72
+ scope; the downloadable release is the publication fallback.
73
+
74
+ ## Deployment limits
75
+
76
+ This is a measured app-testing candidate. Semantic color mistakes remain,
77
+ and smoothing cannot infer an object's unknown original color. Review the
78
+ included failure examples on your app's real input photos before describing
79
+ it as generally production-ready. Browser/mobile performance and real-time
80
+ video have not been validated.
README.md CHANGED
@@ -1,43 +1,80 @@
1
  ---
 
 
2
  tags:
3
- - image-to-image
4
  - colorization
5
  - unet
6
  - pytorch
7
- license: apache-2.0
 
8
  datasets:
9
  - johnowhitaker/imagenette2-320
 
10
  ---
11
- # Mini U-Net Colorizer
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
12
 
13
- A 3,968,892-parameter U-Net that colorizes grayscale photos. Classification-style
14
- (Zhang et al., *Colorful Image Colorization*, [arXiv:1603.08511](https://arxiv.org/abs/1603.08511)): it predicts a distribution over
15
- 236 quantized CIE Lab `a`/`b` bins per pixel rather than regressing a
16
- single ab value directly, with color-bin loss weights derived from the real
17
- training-data distribution (rare/saturated colors weighted higher) so it
18
- doesn't just hedge toward desaturated averages. Decode with an annealed mean.
19
 
20
- - **Status:** Final (20 epochs complete)
21
- - **Input:** L channel normalized as `L/50 - 1` → `[-1, 1]`, shape `(1, 256, 256)`
22
- - **Output:** logits over 236 ab bins, shape `(236, 256, 256)`
23
- - **Trained on:** `johnowhitaker/imagenette2-320` (None), warm-started from `User-2468/mini-unet-colorizer`
24
- - **Loss so far (weighted soft cross-entropy):** train 2.4069, val 2.7111
25
 
26
- ## Usage
 
 
 
27
 
28
  ```python
29
- import numpy as np, torch
30
- from skimage.color import rgb2lab, lab2rgb
31
  from PIL import Image
32
- # paste the SmallUNetColorizer class definition from the training script, then:
33
- model = SmallUNetColorizer.from_pretrained("User-2468/mini-unet-colorizer")
34
- model.eval()
35
- img = Image.open("photo.jpg").convert("RGB").resize((256, 256))
36
- lab = rgb2lab(np.asarray(img).astype("float32") / 255.0)
37
- L = torch.from_numpy(lab[:, :, 0:1] / 50.0 - 1.0).permute(2, 0, 1)[None]
38
- with torch.no_grad():
39
- logits = model(L)
40
- ab = model.decode(logits, temperature=0.38)[0].permute(1, 2, 0).numpy()
41
- L_out = (L[0, 0].numpy() + 1) * 50.0
42
- lab_out = np.concatenate([L_out[:, :, None], ab], axis=-1)
43
- rgb_out = np.clip(lab2rgb(lab_out), 0, 1)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ license: apache-2.0
3
+ pipeline_tag: image-to-image
4
  tags:
 
5
  - colorization
6
  - unet
7
  - pytorch
8
+ - safetensors
9
+ - onnx
10
  datasets:
11
  - johnowhitaker/imagenette2-320
12
+ - detection-datasets/coco
13
  ---
14
+ # Mini U-Net Colorizer — broader-data trained candidate
15
+
16
+ **Status: evaluated app-testing candidate.** Some broad color patches and
17
+ incorrect object hues remain. Predicted colors are not evidence of original
18
+ historical colors.
19
+
20
+ This checkpoint has 3,968,892 learned parameters and 236 fixed color bins.
21
+ It starts from the audited bin-mapping repair of main commit
22
+ `6c47ea40724d8fcd67d4f36ce837dc1cb5b1b2a8` and changes all 65 learned parameter
23
+ tensors. The color vocabulary is unchanged. The selected weights are update
24
+ 748 of a completed 1,122-update BF16 L4 run using a 17,325-photo mixed training
25
+ pool (11,943 Imagenette plus 5,382 COCO), batch 32, initial learning rate 1e-5,
26
+ weighted classification loss and frozen BatchNorm running statistics.
27
+
28
+ Selection compared Imagenette50 and reserved COCO200 validation images.
29
+ The selected model retained color strength better than spatial-loss candidates.
30
+ It was then scored on separate Imagenette200 and COCO-val100 checks.
31
+
32
+ | Test sample | Previous repaired error | This release error | Fine excess-edge reduction |
33
+ |---|---:|---:|---:|
34
+ | Imagenette 200 | 13.318 | 12.847 | 85.8% |
35
+ | COCO-val 100 | 15.266 | 14.532 | 84.3% |
36
 
37
+ Error is mean Lab chroma distance. The release includes guided8 decoding;
38
+ raw learned weights alone improve error by 1.97% and 3.25%, respectively.
39
+ Excess-edge reductions are proxies, not counts of visible blotches removed.
40
+ COCO is a convenience sample; older upstream training exposure is unknown.
 
 
41
 
42
+ ## Use the complete pipeline
 
 
 
 
43
 
44
+ ```bash
45
+ python -m pip install -r requirements.txt
46
+ python inference.py --model . --output-dir colorized photo.jpg
47
+ ```
48
 
49
  ```python
 
 
50
  from PIL import Image
51
+ from model import load_model
52
+ from inference import colorize
53
+ model = load_model('.')
54
+ colorize(model, Image.open('photo.jpg')).save('colorized.png')
55
+ ```
56
+
57
+ Defaults: temperature 0.38, guided radius 8, epsilon 0.001, one network pass.
58
+ The wrapper preserves aspect ratio and original luminance. Old app code that
59
+ only loads safetensors will not automatically gain guided filtering.
60
+
61
+ For ONNX without PyTorch:
62
+
63
+ ```bash
64
+ python -m pip install -r requirements-onnx.txt
65
+ python colorize_onnx.py --model colorizer.onnx --output-dir colorized photo.jpg
66
+ ```
67
+
68
+ The 15.9MB ONNX graph includes the model and guided decoder, with dynamic
69
+ batch/spatial sizes and verified PyTorch parity. See `DEPLOYMENT.md` for the
70
+ Lab input contract, CPU timings, publication commands and integration limits.
71
+ See `RESEARCH_ROUND2.md` and `reports/round2/` for complete measured evidence.
72
+
73
+ ## Limitations
74
+
75
+ Smoothing removes fine color fluctuations but can suppress true small color
76
+ details, especially without luminance boundaries. Semantically wrong hues
77
+ remain. The coffee and rocket failure examples are retained. Browser/mobile
78
+ performance, video consistency and general production quality are unvalidated.
79
+ The separate experiment bundle contains all runs and reproduction code; GPU
80
+ optimizer state is not included. This is a weights-only continuation point.
RESEARCH_REPORT.md ADDED
@@ -0,0 +1,77 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Mini U-Net Colorizer — research and validation, 20 September 2026
2
+
3
+ ## Result and release decision
4
+
5
+ A confirmed bin-mapping bug was repaired. On the reserved 100-image test subset, mean per-image Lab chroma error fell from **31.27 to 13.90 (55.5% lower)**. **92/100** images improved. The model still has **3,968,892 learned parameters**. No learned weights changed in the recommended release; only the 236-by-2 nonlearned color lookup buffer changed.
6
+
7
+ Two actual 100-update fine-tuning runs completed. Neither beat the repaired model on the 50-image validation subset, and both lost colorfulness. Their final trained checkpoints and optimizer states are retained, but they were not selected for release. The current recommendation is the validated decoding repair, not either new trained model.
8
+
9
+ **Production status: improved release candidate, not a general-purpose production model.** The coffee example still shows substantial blotching, and colors can be semantically wrong. The root cause of all remaining artifacts is not established.
10
+
11
+ ## Confirmed failure mechanism
12
+
13
+ The audited main revision was `6c47ea40724d8fcd67d4f36ce837dc1cb5b1b2a8`; stable was `b875f0fd5ff7f2e39f2c6068a60b268fb35586cc`. Both have 236 bins. In both checkpoints, **232 rows of config.json differ from the serialized bin_centers buffer**, with coordinate differences up to 170 Lab units.
14
+
15
+ The original code creates target labels from freshly estimated bins. It then warm-starts all matching tensor shapes, including bin_centers. This can replace the decoder's vocabulary without changing the target encoder. Training therefore optimizes channels in one vocabulary while previews and inference interpret those channels through another. Cross-entropy does not use the decode buffer, so an improving training loss cannot catch the problem.
16
+
17
+ Re-running the original source on the original data, seed-0 split, low-chroma training filter and 3,000-image bin sample recreated **exactly the config vocabulary**: 11,943 filtered training examples, 236 bins. This is evidence for using config for these particular snapshots. It is not a general rule to trust config whenever it conflicts with weights.
18
+
19
+ Integrity verification found only `bin_centers` changed; every learned tensor is equal. Original and revised implementations produced bit-identical raw logits on a divisible-by-eight test input. These findings distinguish a decode repair from retraining or a change to architecture.
20
+
21
+ ## Validation and experiments
22
+
23
+ Validation used 50 images, five per class, from the historical seed-0 5% holdout. All variants use 256-by-256 square preprocessing and temperature 0.38 for comparability. CE is unweighted soft cross-entropy against the same confirmed target bins. Chroma ratio is mean predicted chroma divided by mean reference chroma; proximity to 1 is useful but cannot prove realism. Excess chroma edge is a heuristic for excessive predicted color discontinuities in low-luminance-gradient regions; a gray output can score well, so it must not be optimized alone.
24
+
25
+ | Candidate | CE, lower better | ab error, lower better | Chroma ratio | Excess color edge |
26
+ |---|---:|---:|---:|---:|
27
+ | Existing stable | 2.9584 | 31.40 | 1.895 | 0.907 |
28
+ | Repaired stable | 2.9584 | 14.05 | 1.152 | 0.637 |
29
+ | Existing main | 2.8444 | 30.84 | 1.830 | 0.719 |
30
+ | Repaired main | 2.8444 | 12.77 | 1.021 | 0.520 |
31
+ | Low-context-LR pilot | 2.9927 | 13.20 | 0.616 | 0.351 |
32
+ | High-context-LR pilot | 2.9886 | 13.31 | 0.569 | 0.308 |
33
+
34
+ The two pilots started from repaired main with identical data order, seed, optimizer, augmentation and loss. Each used 100 updates, batch size 2, 256px inputs, a 400-image stratified training pool, backbone/head LR 5e-5, cosine decay, AdamW, gradient clipping and frozen BatchNorm running statistics. The only between-pilot difference was context LR: 5e-5 versus 3e-4. Their priors were re-estimated identically on the fixed checkpoint vocabulary using the small pilot training pool. These settings differ from the historical full-data run; this is **not** a definitive test of longer high-LR training. The candidate-selection criterion was validation CE with visual/colorfulness checks. Neither trained model was promoted.
35
+
36
+ ## Reserved test results
37
+
38
+ The mapping repair was chosen using source reconstruction and the validation subset. A separate 100-image subset, ten per class, was then scored with the same temperature and preprocessing, without using it to tune parameters.
39
+
40
+ | Metric | Existing main | Repaired main |
41
+ |---|---:|---:|
42
+ | Mean per-image ab error | 31.274 | 13.904 |
43
+ | Mean predicted chroma | 27.399 | 14.767 |
44
+ | Reference mean chroma | 15.553 | 15.553 |
45
+ | Chroma ratio | 1.762 | 0.949 |
46
+ | Excess color edge | 0.835 | 0.551 |
47
+ | Soft CE | 2.9750 | 2.9750 |
48
+
49
+ Paired bootstrap over 100 images, 10,000 resamples: mean improvement 17.37 Lab units; conditional 95% interval [15.12, 19.87]. This interval describes this small, stratified benchmark, not broad deployment performance. The equal CE is expected because only decoding changes. Seven test references have low mean chroma; color-source-only metrics are also recorded in JSON.
50
+
51
+ The reconstructed split has no exact-file duplicate leakage among 13,394 examples. Near duplicates and exposure during older upstream training runs have not been ruled out. All split indices and the Parquet SHA-256 are saved. No claims of a pristine never-seen test set across the model's entire history are made.
52
+
53
+ ## Visual and application checks
54
+
55
+ See `reports/branch_bin_comparison.png`, `reports/test_comparison.png`, and `reports/out_of_domain.png`. The comparison removes much of the blue cast and makes many ordinary surfaces more plausible. Four scikit-image examples were used as qualitative probes outside the ten-class benchmark: astronaut, coffee, cat, and rocket. They are far too few to establish out-of-domain reliability. Coffee remains a clear failure case.
56
+
57
+ The production inference wrapper preserves aspect ratio, handles EXIF orientation and odd image dimensions, bounds model input by the longest side, retains original-resolution luminance, and upsamples only chroma. This avoids square stretching and loss of luminance detail. It is a separate preprocessing change from the square-input quantitative benchmark and still needs broader application-specific validation. No GPU/mobile/browser latency claim is made; evaluation timings were collected under concurrent CPU workloads and are unsuitable for performance comparisons.
58
+
59
+ Seven integrity tests pass: sparse versus dense CE and gradients, stable soft encoding, identity context initialization, checkpoint roundtrip and mismatch rejection, shape/temperature/inference handling, and exact reproduction of Hugging Face's split. The original UV trainer is also supplied with guards against repeating the bin mismatch, stable loss reductions, identity initialization and better local saving. The modular fine-tuner is the preferred experiment path because it fixes vocabulary semantics across runs.
60
+
61
+ ## Research implications and next work
62
+
63
+ The original project log overstated the certainty of the receptive-field explanation. It also described the context block as an identity at initialization, but its last BatchNorm scale was initialized to one. New code initializes it to zero; loaded historical context weights are preserved. This audit did not identify whether the remaining blotches originate mainly from semantic uncertainty, limited data, upsampling, or training objectives.
64
+
65
+ [Zhang, Isola and Efros (2016)](https://richzhang.github.io/colorization/) supports classification and rebalancing, but also reports training on roughly one million images and warns about dataset bias. Our inference is that ten Imagenette categories are insufficient evidence for a broad photo product. The paper does not establish that this particular small model must use this formulation forever.
66
+
67
+ [Odena, Dumoulin and Olah (2016)](https://distill.pub/2016/deconv-checkerboard/) motivates testing resize-convolution if periodic artifacts remain. The present 2x2/stride-2 layers do not have uneven overlap; their existence alone does not prove they cause these blotches. Any replacement needs a controlled learned-weight comparison.
68
+
69
+ [DDColor (ICCV 2023)](https://arxiv.org/abs/2212.11613) emphasizes semantic features and multiscale color reasoning. A possible future experiment is teacher-guided training or a compact pretrained semantic encoder, keeping the deployed student under 4M parameters. This is a proposal, not an experiment completed here.
70
+
71
+ Next experiments should first expand diverse, permitted training and evaluation coverage (people, indoor scenes, foliage, food, low-light and historical scans), preserve an untouched external test set, and run a full-data conservative fine-tune against this repaired baseline. Then compare single interventions such as context LR, a coarse chroma output head, or semantic distillation. Choose explicit product targets for acceptable visible artifacts, latency and maximum image size before calling the result production-ready.
72
+
73
+ ## Publication and retained artifacts
74
+
75
+ The connected Hugging Face OAuth credential allowed Jobs/read access but lacked write-repos. A small direct-main upload failed with HTTP 403; the suggested PR route also failed with 403. No Hub branch was modified. The later user-selected `Huggingface` custom plugin did not expose separately callable tools in this session; its actual permissions could not be tested.
76
+
77
+ `release/` is ready for an atomic main-branch upload via `upload_main.py` after normal write authentication. The complete trained checkpoints, logs, scripts, split manifests, metrics and recommended repair are included in the experiment bundle. Stable remains at the original revision. No training job remains running.
RESEARCH_ROUND2.md ADDED
@@ -0,0 +1,234 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Mini U-Net Colorizer: spatial artifacts and further training
2
+
3
+ 20 September 2026. This study starts from the **bin-mapping repair**, not the
4
+ broken Hub checkpoint. Improvements below must not be conflated with the
5
+ previous repair's 55.5% reduction in chroma error.
6
+
7
+ ## Selected release and final results
8
+
9
+ **Selected:** mixed-data GPU checkpoint at update 748, from the completed
10
+ 1,122-update run. Selection used 50 Imagenette plus 200 reserved COCO training
11
+ images, with equal weight for each domain's relative chroma error. A candidate
12
+ had to retain at least 90% of baseline predicted chroma in both domains and
13
+ avoid worse fine excess-edge scores. CPU and GPU spatial-loss candidates
14
+ failed the COCO color-retention gate (87.0% and 89.0%). The selected checkpoint
15
+ retained 96.5% and 97.9%, respectively, with a 2.66% improvement in the
16
+ selection score. No final-test results were used to change weights or decoding.
17
+
18
+ All comparisons below start from the already repaired baseline. Lower error
19
+ and excess-edge scores are better. Chroma ratio is predicted/reference mean
20
+ chroma; it is not a realism score.
21
+
22
+ | Test sample | Pipeline | Chroma error | Chroma ratio | Fine excess edge | Coarse excess edge |
23
+ |---|---|---:|---:|---:|---:|
24
+ | Imagenette 200 | Previous repair, raw | 13.318 | 0.965 | 0.5395 | 1.1962 |
25
+ | Imagenette 200 | Previous repair, guided8 | 13.077 | 0.943 | 0.0829 | 0.5948 |
26
+ | Imagenette 200 | New weights, raw | 13.056 | 0.919 | 0.5188 | 1.0782 |
27
+ | Imagenette 200 | **New weights, guided8** | **12.847** | **0.899** | **0.0765** | **0.5354** |
28
+ | COCO-val 100 | Previous repair, raw | 15.266 | 0.981 | 0.5896 | 1.4608 |
29
+ | COCO-val 100 | Previous repair, guided8 | 14.992 | 0.959 | 0.0990 | 0.7479 |
30
+ | COCO-val 100 | New weights, raw | 14.771 | 0.985 | 0.5813 | 1.3267 |
31
+ | COCO-val 100 | **New weights, guided8** | **14.532** | **0.966** | **0.0923** | **0.6727** |
32
+
33
+ Relative to the previous raw repair, the complete pipeline improves chroma
34
+ error by **3.54% and 4.81%**, fine excess edges by **85.83% and 84.35%**, and
35
+ coarse excess edges by **55.24% and 53.95%**, on Imagenette and COCO respectively.
36
+ These are measured proxies, not percentages of visible blotches eliminated.
37
+ Mean error improves on 142/200 and 76/100 images. Paired-image bootstrap 95%
38
+ intervals for mean error reduction are [0.308, 0.639] and [0.362, 1.089] Lab units.
39
+
40
+ There is a learned improvement beyond postprocessing: raw weights improve
41
+ error by 1.97% and 3.25%; with the same guided8 decoder on both old and new
42
+ weights, improvement is 1.76% and 3.07%. The corresponding guided-to-guided
43
+ paired intervals are [0.060, 0.398] and [0.091, 0.812]. Per-image results,
44
+ 10,000-resample bootstrap calculations and low-chroma-source summaries are
45
+ provided in `reports/round2/final_results.json` and the final evaluation JSONs.
46
+
47
+ Qualitative checks retain failures. The coffee image has substantially fewer
48
+ fine rainbow patches but still contains wrong broad hues. The cat remains
49
+ undercolored. The rocket's color error worsens from 21.94 to 22.86 despite
50
+ smoothing. Four fixed probes and one fixed example per Imagenette class are
51
+ shown in `final_visual.png` and `final_test_comparison.png`; they were not
52
+ selected for favorable outcomes. The held-out quantitative pipeline resizes
53
+ images to 256x256; the four application probes exercise the aspect-preserving,
54
+ full-resolution output wrapper. Their metrics are not directly pooled.
55
+
56
+ The result is a useful **app-testing candidate**, not evidence that all
57
+ blotches are removed or that arbitrary real-world photos are production-ready.
58
+
59
+ ## Deployment verification
60
+
61
+ The selected safetensors file has SHA-256
62
+ `0e4c417375684a044860f8af3ac3a2fb44e1a5729254ca33abda758ea71aea6e`.
63
+ All 65 learned parameter tensors changed from the repaired baseline; the
64
+ verified color-bin buffer is identical. Parameter count remains 3,968,892.
65
+
66
+ An ONNX opset-17 graph includes the learned model, temperature-0.38 decoding
67
+ and guided8 filtering. It is 15,885,757 bytes and accepts dynamic spatial sizes
68
+ and batches. Four numerical shape checks passed, with maximum Lab-ab difference
69
+ below 0.000009 against PyTorch. Three image-wrapper comparisons preserved image
70
+ dimensions and differed by at most one 8-bit RGB level. Ten code tests passed.
71
+ The wrapper and graph input/output contract are documented in `DEPLOYMENT.md`.
72
+
73
+ On this shared CPU host, ONNX Runtime 1.30.0 with two threads measured median
74
+ 909ms and p90 1239ms for a 256x256 input (20 iterations after five warmups).
75
+ Peak process RSS was 508MiB. This includes network, decoding and filtering,
76
+ excludes file loading/full-resolution color conversion, and is not a browser
77
+ or mobile latency guarantee. The graph is not quantized.
78
+
79
+ Repository publication remains blocked: only the original Hugging Face
80
+ connection is callable, with Jobs/read scopes and no write-repos scope.
81
+ Previous direct and PR upload attempts returned 403; rechecking scopes at
82
+ completion confirmed the same restriction. Neither main nor stable changed.
83
+ The complete release, source, evidence and atomic main-upload helper are
84
+ preserved for manual publication. No additional training is left running.
85
+
86
+ ## Questions and evidence
87
+
88
+ Three interventions were tested: spatial decoding, supervised spatial loss,
89
+ and broader training data. Five training runs completed, including three L4
90
+ GPU runs. Every model retains 3,968,892 learned parameters and the same 236
91
+ color-bin meanings. Model selection uses validation images, with separate
92
+ Imagenette and COCO external checks. Final quantitative results and the selected release follow below.
93
+
94
+ ### Spatial decoding
95
+
96
+ Thirteen alternatives were compared on 50 validation images: the repaired
97
+ baseline; temperatures 0.6 and 0.8; logit pooling by factors 2, 4 and 8;
98
+ luminance-guided filtering with radii 4, 8 and 16; horizontal-flip averaging;
99
+ flip averaging plus radius-4 filtering; pooling plus filtering; and a gray
100
+ control. The gray control achieves zero discontinuity scores, demonstrating
101
+ why these scores cannot select a model by themselves.
102
+
103
+ The [guided-filter paper](https://people.csail.mit.edu/kaiming/eccv10/index.html)
104
+ motivates a local linear filter whose coefficients depend on luminance. This
105
+ implementation filters predicted Lab chroma using only input luminance;
106
+ reference colors are never passed into inference. It can smooth spurious
107
+ color changes while retaining luminance boundaries. Same-luminance color
108
+ boundaries remain a limitation.
109
+
110
+ Radius 8 was nominated as the balanced single-pass mode using validation
111
+ results and images. Radius 16 visibly removes more genuine small color
112
+ details, despite better artifact scores. Flip averaging plus radius 4 is an
113
+ optional two-pass mode. Temperature stays 0.38 and saturation stays 1.0.
114
+
115
+ On the additional 200-image Imagenette test, radius-8 filtering alone changed
116
+ mean chroma error from 13.318 to 13.077, fine excess color edge from 0.5395 to
117
+ 0.0829, and coarse excess edge from 1.1962 to 0.5948. Mean predicted chroma
118
+ relative to reference changed from 0.965 to 0.943. Error improved on 179/200
119
+ images; paired image bootstrap 95% interval for the mean improvement was
120
+ [0.205, 0.276] Lab units.
121
+
122
+ On 100 COCO-val photos, filtering alone changed error from 15.266 to 14.992,
123
+ fine excess edge from 0.5896 to 0.0990, and coarse excess edge from 1.4608 to
124
+ 0.7479. Chroma ratio changed from 0.981 to 0.959. Error improved on 95/100
125
+ images; the corresponding interval was [0.233, 0.314]. These intervals describe
126
+ the image samples, not performance across arbitrary deployment photos.
127
+
128
+ ### Learned spatial consistency
129
+
130
+ The added loss matches chroma gradients to reference gradients at scales
131
+ 1, 4 and 16. Chroma is normalized by 110; the gradient differences use smooth
132
+ L1 with beta 0.05, averaged over scales and horizontal/vertical directions.
133
+ Its coefficient is 10 alongside the original weighted classification loss.
134
+ This is supervised transition matching, rather than a requirement that every
135
+ region become uniformly gray.
136
+
137
+ The CPU pair used 400 updates, batch 2, learning rate 2e-6 and seed 0. The
138
+ GPU pair used 1,122 updates, batch 32, learning rate 1e-5 and seed 123: three
139
+ complete passes through 11,943 training photos. Within each pair, the only
140
+ training-objective difference is the spatial loss. Both use frozen BatchNorm
141
+ statistics, identical data order, fixed audited color weights, and cosine
142
+ learning-rate decay. CPU-to-GPU differences also change batch, precision,
143
+ seed and training duration and are not a single-factor comparison.
144
+
145
+ At the final CPU checkpoint, classification-only versus spatial loss produced
146
+ validation error 12.384 versus 12.149, CE 2.8076 versus 2.7918, and chroma
147
+ ratio 0.962 versus 0.923. At the final GPU checkpoint, error was 12.735 versus
148
+ 12.326 and CE 2.8293 versus 2.8091. Earlier CPU checkpoints had lower error
149
+ but appreciably less color, illustrating the need to inspect colorfulness.
150
+
151
+ These results support the added loss in this experiment; they do not prove
152
+ that the existing artifacts have one particular architectural cause.
153
+
154
+ ### Broader data
155
+
156
+ A third GPU arm added 5,382 eligible COCO training photos to the 11,943
157
+ Imagenette photos. The first two pinned COCO training shards contain 5,864
158
+ photos; 200 were reserved for validation and the remaining photos were
159
+ filtered using the same low-chroma threshold. The combined training pool has
160
+ 17,325 photos. The run used the classification objective and the same 1,122
161
+ updates and batch size as the GPU control, approximately 2.07 epochs over
162
+ the larger pool. The original color vocabulary and class weights stayed fixed.
163
+
164
+ The COCO training-validation set is distinct from the 100 COCO-val photos
165
+ used as an external probe. Image-ID intersection between the training shards
166
+ and that external sample is zero. The COCO revision is
167
+ `26ddc382fe75dfc2a0655b5977e296ea10efebce`. The actual indices, hashes and
168
+ training manifests are retained. Broader-data training was exploratory;
169
+ this is not a claim that a small COCO subset is sufficient for production.
170
+
171
+ ## What the diagnostics establish
172
+
173
+ The constant-luminance and shifted-image probes show spatial sensitivity.
174
+ Across ten validation images, shifting by one pixel changed predicted chroma
175
+ by 2.31 Lab units on average in the central region; an eight-pixel shift
176
+ changed it by 1.45. Flat-input responses also contain spatial color variation.
177
+ This does not isolate a causal layer: padding, pooling, dilations, upsampling
178
+ and learned features can all contribute.
179
+
180
+ [Odena et al.](https://distill.pub/2016/deconv-checkerboard/) motivates testing
181
+ upsampling effects, but the current kernel-2/stride-2 layers do not have the
182
+ classic uneven-overlap configuration. Replacing them without a controlled
183
+ training comparison is not established here as a fix.
184
+
185
+ The current context path already covers hundreds of input pixels; simply
186
+ increasing dilation is not a demonstrated remedy. The remaining wrong hues
187
+ also point to a semantic problem that spatial filtering cannot solve. The
188
+ [original colorization research](https://richzhang.github.io/colorization/)
189
+ discusses dataset bias despite much larger training coverage, and
190
+ [DDColor](https://arxiv.org/abs/2212.11613) motivates semantic and multiscale
191
+ features. Semantic teacher-to-student training or a pretrained compact
192
+ encoder remain untested architectural alternatives here, not completed work.
193
+
194
+ ## Evaluation limits
195
+
196
+ The 200-image Imagenette test is disjoint from the previous 100-image test
197
+ and the 50-image validation set. It comes from the historical seed-0 5%
198
+ holdout. Exposure during older upstream training remains unknown. The 100
199
+ COCO-val images are the first viewer rows, a convenience sample with 72
200
+ object categories represented, not a population-representative random test.
201
+
202
+ Chroma error measures closeness to one reference colorization; plausible
203
+ alternative colors may score poorly. Excess-edge measures are heuristics,
204
+ not counts of visible blotches or human quality ratings. Grayscale-source
205
+ references are separately counted and color-source-only summaries retained.
206
+ Qualitative examples include known failures, not only favorable results.
207
+
208
+ The additional 200-photo COCO training-validation comparison uses the reserved
209
+ indices, with 256px JPEG95 copies for transfer. Every candidate receives the
210
+ same decoded copies. These images are used for selecting weights and are
211
+ not reported as an independent final test.
212
+
213
+ ## Execution and reproducibility
214
+
215
+ The custom plugin named Huggingface still exposed no tools. The original
216
+ Hugging Face connection supplied Jobs and read access but lacked write-repos.
217
+ A 17 MiB checksum-verified log roundtrip established a fallback artifact
218
+ transport. Larger two-checkpoint archives exceeded the tool's 64 MiB response
219
+ frame because the response repeats log content; CPU helpers exported one
220
+ checkpoint at a time. Explicit log-tail limits avoided server truncation.
221
+ All recovered archives are checked by size, SHA-256 and ZIP CRC.
222
+
223
+ Two initial GPU starts stopped on nonfinite FP16 gradients before completing
224
+ training. BF16 restarts completed with finite-gradient guards enabled. Those
225
+ failed starts are recorded separately from the five completed training runs.
226
+ The three GPU jobs were `6aaf8ceb52d0dbd7f1d73094`,
227
+ `6aaf8cf752d0dbd7f1d7309a`, and `6aaf8dad51992417dfccc557`.
228
+
229
+ Both best and last learned checkpoints are retained. GPU optimizer state was
230
+ not exported; those checkpoints support a new-optimizer fine-tune, not exact
231
+ training resume. CPU optimizer states are retained. Source, configurations,
232
+ environments, priors, split definitions and per-image results accompany them.
233
+ The corrected decoder and vocabulary guards continue to prevent the earlier
234
+ silent mapping mismatch.
SHA256SUMS.json ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "DEPLOYMENT.md": "089cadec89caa645a5f2617f869ec50186dd8db574b0680a51d4bc12f84b96cb",
3
+ "README.md": "f05ab35140319d6d8d5a65669d62645e3909b1ec1bd233bd66e0af3c1b135eb3",
4
+ "RESEARCH_REPORT.md": "6730cf20f7a584707fb0c084b882ac87db1ef858d72e7d5e716e5e0d382ddca1",
5
+ "RESEARCH_ROUND2.md": "19f42c72003cb685f46b315105ad2d4e62a10a7e0cd1455ab11db51bd5083728",
6
+ "colorization_project_log.md": "66dbd61e5e45f208fdd36525f19c232678892295d6b31882a5fb71627cb731b9",
7
+ "colorize_onnx.py": "4ec6bc9dba4a998243e97b3bd824d25639455e6246a47886be1ee796121cf1ba",
8
+ "colorizer.json": "9f1db0072e91747bd6e9d6de1e2387440a4a2e0669554e499bbf29288a9758cd",
9
+ "colorizer.onnx": "0ef86749901e66ad53b1e8e1d940330572e2f4c9347ac01e7bc02c7683f8c79a",
10
+ "config.json": "90299ee25ee3c3e645b362ddbd5a859b9e739c2a0a1c8a5e9ea79df171288bff",
11
+ "export_onnx.py": "4b0b8494f0eba3e813457fe393c3b96ce1ab35fefcf5f63a9251f4e852552737",
12
+ "inference.py": "db958511d7a9eecec4c5730d96ac3f0005e002720b1a7f4b035e29cac7e7c5e3",
13
+ "model.py": "4cc57f82ffd6378bdf23a088ce7a1ed56c09e5673de3ee5232b8ee1cacc9be0a",
14
+ "model.safetensors": "0e4c417375684a044860f8af3ac3a2fb44e1a5729254ca33abda758ea71aea6e",
15
+ "reports/round2/coco_manifest.json": "80061d796a1e4f08b508203cd45513916f2f0e74efa0c516673413c44f9d6dcb",
16
+ "reports/round2/coco_revision.json": "c83949c5adf6e0e29dfb1d37c6beff0044bfa6a6a7161a6e0264bdbf41afd450",
17
+ "reports/round2/coco_validation200.json": "bf6adcb05a2a0f92a4f428073ff95ff9977064a2a99b3b911ee5c3a5797955ea",
18
+ "reports/round2/environment-lock.txt": "f95748796cbc5d0667fe17163b6054fb70567a9e87206473e743d150adc26e45",
19
+ "reports/round2/final_coco100.json": "dabe07889d4ce4fc512f60768aa5c3c6268060464e21aed9c4c1fe28294675a1",
20
+ "reports/round2/final_coco100.png": "2ba6c18c60e7858574f45f1fe6b1d25ef14031d8b92a3986ec057f5c945c337e",
21
+ "reports/round2/final_integrity.json": "de32a8c9173efdc112c776a4937a5024b61814e93b16818eeccf2c43bc60b62a",
22
+ "reports/round2/final_results.json": "40e2e317c3d083be9dedefdcc0ce942d0b972fd64164201680a6aeab589a870b",
23
+ "reports/round2/final_test200.json": "d525d24e8706ff3f7ee3478c1ab95b8d5875d0d8242f6389dc29b9416449ce63",
24
+ "reports/round2/final_test_comparison.png": "9681e45700aa5def0f69f67a49a45cadfddb06c1dc4b321995e09641eebf7f73",
25
+ "reports/round2/final_visual.json": "0ca89a7d8017046c085c7983374706df0c1970284d95a8519cf16bf4d65065a8",
26
+ "reports/round2/final_visual.png": "df2b3c271150059ab2f0d0a10f7fe889f1d3878c6096e14e0fd5df13edd7022a",
27
+ "reports/round2/hub_refs.json": "e714d59b825a333abbea48270353a8dffd5c1b1b8b629028a729fac063f7cbd8",
28
+ "reports/round2/jobs.json": "aa8dc41c57ae6ee41b5e51e8aff09055dafb536424413240cbeff3fcbabcbacc",
29
+ "reports/round2/manifest.json": "adaf3a771a4c1cd3e1c58e6d74eae2a164e1d00ddbca8232067711111caa62ea",
30
+ "reports/round2/onnx_benchmark.json": "2afbcbd29f9619dcc0cad1954f8fe2af3fdd11b94235b53883e8e0db53ef67f0",
31
+ "reports/round2/onnx_export.json": "9f1db0072e91747bd6e9d6de1e2387440a4a2e0669554e499bbf29288a9758cd",
32
+ "reports/round2/selection_protocol.md": "c297f6141a6c5419c2d6bba4ea1ecb5ce7295a78f9d16abec9e38224d5000a7b",
33
+ "reports/round2/tests-final.txt": "60e907edbabdfcc838d135fbd18b7733ab87c401bbf608c734f589abcc4687f8",
34
+ "reports/round2/weight_selection.json": "5ca680f16fab8fedb1ca1a14f59a127ee5d5a71e21fb3560f78417b530201df0",
35
+ "requirements-onnx.txt": "c243944cabecd501038b8027dc5fd08f6d2ce74bcc47cc121f5c8df1b8c4445c",
36
+ "requirements.txt": "a5b6a1804d983178099da00bad2f92a4196009384e431119cc73c5433d249622",
37
+ "spatial.py": "d13abe399ef049e21a6459a7003461afaae0c00e0c560262c6cf375de4c9884a",
38
+ "training/color_statistics.npz": "1d14920885324137169e63f908a413ebf083e8db1e604d2333847fd2094f36a6",
39
+ "training/environment.txt": "ee1e54f3b86a977f74fbc0283ce0cfd2fcf54cc4c8ec4b9096634746eadcc18b",
40
+ "training/history.json": "714ff0c8364e33e75cff21b6abe94f78db6c4536faf016ad1f591eb0ebb7dc72",
41
+ "training/run_config.json": "6cdd4da9e0ae293e8f1895dd03d28bfb1f45c3abf19fc147cd8c18e6dfd36559",
42
+ "training/split_manifest.json": "dc5901733751e8585f959918e46ce75b520c6e241a79d26ab9b3df943a12ac38",
43
+ "upload_main.py": "d59fe5ebfc27f6051b7e7c6d8ad66b23092fa61b3db9fc87b33bff0fad2b483a"
44
+ }
colorization_project_log.md ADDED
@@ -0,0 +1,251 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ > Current verified findings are in sections 12–13; section 13 is the latest release. Earlier sections are historical and include assumptions overturned by the bin-mapping audit.
2
+
3
+ # Mini U-Net Photo Colorizer — Project Log
4
+
5
+ **Model repo:** `User-2468/mini-unet-colorizer`
6
+ **Scripts:** `colorize_train.py`, `colorize_image.py`, `colorize_eval.py` (self-contained UV scripts, run via `hf jobs uv run`)
7
+ **Infrastructure:** Hugging Face Jobs (CLI-only management: `hf jobs ps`, `hf jobs logs <job_id>`)
8
+
9
+ ---
10
+
11
+ ## 1. Project Goal
12
+
13
+ Train a compact (<4M parameter) U-Net that takes a grayscale (L-channel) photo and predicts plausible, vivid color (Lab `a`/`b` channels), without collapsing to the desaturated "safe average" that naive regression approaches produce. The project has gone through two major architectural pivots and is currently in a third round of tuning aimed at a specific, persistent visual artifact.
14
+
15
+ ---
16
+
17
+ ## 2. Architectural Evolution
18
+
19
+ ### Phase 1 — Regression baseline (abandoned)
20
+ - Predicted `ab` directly from `L` using a `SmoothL1Loss`.
21
+ - **Failure mode:** systematic desaturation. This is the classic regression-to-the-mean problem in colorization: when several colors are equally plausible for a given gray region, a regression loss is minimized by predicting their average, which is close to gray. Confirmed as the root cause and motivated a full rebuild rather than further tuning of the regression head.
22
+
23
+ ### Phase 2 — Classification-style rebuild (current backbone)
24
+ - Rebuilt around the approach from Zhang et al., *"Colorful Image Colorization"*:
25
+ - Quantize Lab `ab` space into a discrete set of bins.
26
+ - Predict a per-pixel **distribution over bins** (soft cross-entropy against the 5 nearest bins per pixel) instead of a single continuous value.
27
+ - Apply **class-rebalancing weights**, computed empirically from the real training-color distribution, so rare/saturated colors aren't drowned out by the abundance of near-neutral pixels.
28
+ - Decode using an **annealed mean**: a temperature-controlled interpolation between the full expectation (smooth, but re-introduces the regression-style hedging) and the argmax/mode (vivid, but can be blotchy).
29
+ - This eliminated the systematic desaturation problem.
30
+
31
+ ### Phase 3 — Dilated-convolution context block (current, still being tuned)
32
+ - Added `DilatedContextBlock` at the bottleneck (see §3.2) to expand the effective receptive field, specifically to address a recurring artifact (see §5).
33
+ - Uses a residual connection so it starts as a near no-op and can be warm-started onto a checkpoint that predates the block without disrupting existing learned behavior.
34
+
35
+ ---
36
+
37
+ ## 3. Current Architecture (`SmallUNetColorizer`)
38
+
39
+ Defined identically (and independently, for Jobs self-containment) in all three scripts.
40
+
41
+ ### 3.1 Backbone
42
+ Standard 4-level U-Net:
43
+ - `enc1`–`enc4`: `double_conv` blocks (Conv3x3 → BN → ReLU ×2), channel widths `base, base×2, base×4, base×8` (default `base=44`).
44
+ - `MaxPool2d(2)` between encoder stages.
45
+ - Symmetric decoder (`dec1`–`dec3`) with `ConvTranspose2d` upsampling and skip connections concatenated from the corresponding encoder stage.
46
+ - `out_conv`: 1×1 conv from `base` channels to `num_bins` raw logits (no activation).
47
+ - Total parameters: ~3.6M at `base=44` (`bin_centers` is a non-trainable registered buffer, not counted against the params budget, which is asserted `< 4,000,000` in the training script).
48
+
49
+ ### 3.2 `DilatedContextBlock` (the Phase 3 addition)
50
+ Sits at the bottleneck, applied to the output of `enc4`, before decoding starts.
51
+
52
+ - `proj_in`: 1×1 conv down to a narrow `mid_ch` (default 96).
53
+ - `dilated`: a stack of 3×3 convs at dilations `(2, 4, 8)` by default — same parameter cost as ordinary 3×3 convs, but much larger spatial reach.
54
+ - `proj_out`: 1×1 conv back up to the original channel width, batch-normed, **no activation before the residual add**.
55
+ - Residual connection + ReLU: `relu(x + y)`.
56
+
57
+ **Design rationale (from code comments):**
58
+ - The plain U-Net's receptive field is fixed in absolute pixels (~51–54px, empirically measured) regardless of input resolution — at larger `--image-size` this covers a shrinking fraction of the image, which was hypothesized as the cause of color decisions that don't stay consistent across a whole object (e.g. the parachute-canopy and chain-saw artifacts, see §5).
59
+ - A naive fix (another U-Net level with channel doubling) would cost ~4.5M params on its own — over budget.
60
+ - The dilated block, in isolation, was empirically measured (via gradient measurement) to reach ~232×232px receptive field with dilations `(2,4,8)`.
61
+ - The residual formulation means it starts as a small perturbation on top of an already-trained backbone, growing in influence as training proceeds — important for warm-starting from a pre-context-block checkpoint.
62
+
63
+ ### 3.3 Decode step
64
+ ```python
65
+ def decode(self, logits, temperature=0.38):
66
+ logp = F.log_softmax(logits, dim=1)
67
+ probs_t = F.softmax(logp / temperature, dim=1)
68
+ return torch.einsum("bqhw,qc->bchw", probs_t, self.bin_centers)
69
+ ```
70
+ - `temperature → 1`: full expectation (smooth, can desaturate).
71
+ - `temperature → 0`: approaches the mode (vivid, can be blotchy/noisy per-pixel).
72
+ - Default `0.38` follows Zhang et al.
73
+
74
+ ---
75
+
76
+ ## 4. Training Pipeline (`colorize_train.py`)
77
+
78
+ ### 4.1 Color bin construction (`ColorBins` / `build_color_bins`)
79
+ - Samples `--bin-samples` (default 3000) training images, resizes to 64×64, converts to Lab.
80
+ - Histograms `ab` values on a `--bin-grid-size`-spaced grid (default 10, range ±110) — a data-driven stand-in for computing exact sRGB gamut geometry; bins that never occur in the sample simply aren't included.
81
+ - **Grayscale-source exclusion from statistics:** images whose mean Lab chroma (`sqrt(a²+b²)`, averaged over pixels) is below `--grayscale-chroma-threshold` (default 3.0) are excluded from the histogram, because a genuinely black-and-white source photo's correct `ab` really is ~0 — a different phenomenon from the model hedging toward gray on a photo that actually had color. Mixing the two would miscalibrate rebalancing.
82
+ - Bin counts are Gaussian-smoothed (`smoothing_sigma`), then converted to rebalancing weights:
83
+ ```
84
+ w = 1 / ((1 - λ) * prior + λ / Q)
85
+ ```
86
+ normalized so `E_prior[w] == 1`. `λ` (`--rebalance-lambda`, default 0.5) trades off pure inverse-frequency weighting (`λ=0`, aggressive) against uniform/no rebalancing (`λ=1`).
87
+ - Soft-encoding (`ColorBins.soft_encode`): for each pixel's continuous `ab`, finds the `k=5` nearest bin centers (via `cKDTree`) and computes Gaussian-kernel soft weights (`sigma=5.0`) over them, normalized to sum to 1 — the per-pixel training target.
88
+
89
+ ### 4.2 Loss (`classification_loss`)
90
+ Weighted soft multinomial cross-entropy:
91
+ - Gathers logits/log-probabilities only at each pixel's 5 nearest bins (not the full `(B, Q, H, W)` tensor) — this is a deliberate memory optimization, since with `Q` in the hundreds this would otherwise be the single largest activation in the model, larger than anything in the backbone; this matters for fitting on a 16GB GPU.
92
+ - Per-pixel weighted NLL is scaled by the rebalancing weight of that pixel's single nearest bin.
93
+
94
+ ### 4.3 Learning-rate groups
95
+ Three separate AdamW parameter groups:
96
+ | Group | Params | Flag | Default |
97
+ |---|---|---|---|
98
+ | Backbone | everything except `context.*` and `out_conv.*` | `--lr` | 2e-4 |
99
+ | Context block | `context.*` | `--context-lr` | falls back to `--lr` |
100
+ | Output head | `out_conv.*` | `--head-lr` | falls back to `--lr` |
101
+
102
+ Rationale for separate LRs: when warm-starting (`--init-from`) onto a checkpoint whose architecture predates a component (context block) or whose output shape differs (head, since it depends on the bin grid), that component starts randomly initialized and may not receive enough signal at the backbone's cautious fine-tuning rate to develop meaningful behavior within a typical epoch budget.
103
+
104
+ Cosine annealing LR schedule (`CosineAnnealingLR`, `T_max = epochs`) applied across all groups.
105
+
106
+ ### 4.4 Warm-starting (`--init-from`)
107
+ - Loads **raw tensors only** (`load_raw_state_dict`) rather than instantiating the source model class — necessary because a checkpoint may come from an earlier, architecturally incompatible version (e.g. the old regression head), whose `config.json` couldn't even construct the current model class.
108
+ - Only tensors whose name **and shape** match the current model are loaded (`strict=False`); everything else is left randomly initialized and reported explicitly in logs.
109
+
110
+ ### 4.5 Dataset loading (`get_dataset_splits`)
111
+ - Handles both directly-loadable datasets and legacy-loading-script datasets (like `frgfm/imagenette`) by falling back to loading the auto-converted Parquet mirror directly from `refs/convert/parquet/<config>/`.
112
+ - If a dataset has no `validation`/`test` split (e.g. `johnowhitaker/imagenette2-320`), a reproducible 5% held-out split is carved out via `.train_test_split(test_size=0.05, seed=args.seed)`.
113
+ - `--exclude-grayscale-source` (off by default): optionally drops grayscale-source images from the actual train/val sets (not just the bin-statistics sample), since they contribute pure "predict zero" gradient that works against the rebalancing objective. Framed as worth trying, not yet a default.
114
+
115
+ ### 4.6 Sample-grid preview fix
116
+ - The recurring training-time sample grid used to draw from the literal first batch of an unshuffled val loader — for a class-grouped dataset like Imagenette, this meant every preview sample came from a **single class**, corrupting the qualitative evaluation signal for the whole project until caught.
117
+ - Fixed by seeding a `np.random.default_rng` and drawing a fixed random (but reproducible) sample spanning the full validation set, held constant across the run.
118
+
119
+ ### 4.7 Checkpointing
120
+ - Pushes to the Hub every `--push-every` epochs (default 10) during training (not just at the end), so a Jobs timeout only loses progress since the last push.
121
+ - Auto-generates a model card (`build_model_card`) with current train/val loss, dataset, warm-start source, and status (in-progress checkpoint vs. final).
122
+
123
+ ---
124
+
125
+ ## 5. Evaluation (`colorize_eval.py`)
126
+
127
+ Runs CPU-only (`--flavor cpu-basic`), since it's inference-only on a ~3.6M param model.
128
+
129
+ - **Class-stratified sampling:** samples `--per-class` images from *every* class (not just a fixed grid), avoiding the single-class bug that once affected the training preview.
130
+ - Maps WordNet synset IDs (used by some Imagenette repo variants, e.g. `johnowhitaker/imagenette2-320`) to readable class names via a hardcoded `SYNSET_TO_NAME` table, so grid labels stay consistent regardless of which repo's labeling convention is active.
131
+ - For each sampled image, produces a (gray | prediction | ground truth) row, and separately flags images whose ground truth is itself grayscale-source (mean chroma below threshold) — tagged in red in the output grid — so a "weak" result there isn't misread as a model failure (there was nothing to recover).
132
+ - Since a Jobs container is destroyed on completion, results are only retrievable via `--push-to-hub` (grid image pushed directly to the model repo).
133
+
134
+ ---
135
+
136
+ ## 6. The Persistent Artifact Problem
137
+
138
+ **Symptom:** hard-edged blocks of inconsistent color within what should be a single coherently-colored object or region — most visibly on a **parachute canopy** and in a **chain-saw scene**. The same failure pattern has recurred across four separate interventions:
139
+
140
+ 1. **Learning-rate tuning** — did not resolve it.
141
+ 2. **Rebalancing lambda adjustment** (`--rebalance-lambda`) — did not resolve it.
142
+ 3. **Decode temperature tuning** (`--temperature` / `--annealed-temp`) — did not resolve it.
143
+ 4. **Receptive-field expansion via the dilated context block** — added most recently, but the run that added it trained the newly-initialized context branch **at the backbone's cautious learning rate (5e-5)**, which is likely too low for a freshly-initialized residual branch to develop meaningful behavior within the run's epoch budget. This confounds the result: the context block's *idea* has not yet been properly tested, only tested under a probably-inadequate LR.
144
+
145
+ **Status:** open question. It is not yet established whether the artifact is a receptive-field limitation, a rebalancing-calibration issue, or something else — the dataset's grayscale-source confound (see below) is a plausible contributing factor to rebalancing miscalibration, but nothing has been confirmed as the true cause pending the next, properly-configured run.
146
+
147
+ ---
148
+
149
+ ## 7. Dataset Notes
150
+
151
+ - **Active dataset:** `johnowhitaker/imagenette2-320` — confirmed healthy, stored natively as Parquet, no legacy-loading-script issues. No validation split (uses `train_test_split` fallback, see §4.5).
152
+ - **Avoid:** `frgfm/imagenette` — intermittently returns 500 errors from `hub_repo_details` on its Parquet mirror; unreliable enough to have been dropped as the default in both the training and eval scripts, though it remains supported as a `--dataset-config` option (`160px`/`320px`/`full_size`) if the mirror recovers.
153
+ - **Confirmed confound:** a meaningful fraction of the training images are themselves originally grayscale (not just desaturated-but-color). These images contribute a "correct" `ab ≈ 0` gradient, which is statistically indistinguishable from the model incorrectly hedging toward gray, and this complicates calibration of the rebalancing weights (`build_color_bins` already excludes them from the *bin-derivation statistics*, but by default they remain in the actual training set — the `--exclude-grayscale-source` flag exists to test removing them entirely, and hasn't yet been adopted as default).
154
+ - **Rejected fallback:** using the `160px` Imagenette config was considered as a way to sidestep dataset issues, but rejected — upsampling 160px source images back up to the training resolution would undermine the entire point of training at `320px` native resolution.
155
+
156
+ ---
157
+
158
+ ## 8. Key Learnings So Far
159
+
160
+ 1. **Regression → classification is a hard requirement, not a tuning knob.** Direct `ab` regression collapses to desaturated averages regardless of loss-weighting tricks; only a rebalanced classification formulation with a genuinely multi-modal target produced vivid output.
161
+ 2. **Grayscale-original images are a first-class confound.** They must be treated separately from "the model desaturated a color photo," both when calibrating rebalancing statistics and, potentially, in what data the model trains on at all.
162
+ 3. **Newly-initialized branches need their own learning rate.** A branch initialized from scratch (context block, or a classification head grafted onto a warm-started backbone) does not reliably develop behavior if it only ever sees the backbone's cautious fine-tuning LR — this pattern has shown up twice now (head LR, then context LR).
163
+ 4. **Diagnostic tooling bugs can silently corrupt the evaluation signal.** The single-class sample-grid bug (unshuffled first-batch selection) was invisible unless someone happened to notice every preview image belonged to the same object category.
164
+ 5. **Upsampling defeats the purpose of a resolution choice.** Don't fall back to a lower-resolution dataset config as a workaround for unrelated pipeline issues if it requires upsampling back to the target training resolution.
165
+
166
+ ---
167
+
168
+ ## 9. Current State / Most Recent Run
169
+
170
+ - Context block added to the architecture.
171
+ - Trained with the newly-initialized context branch at the backbone's LR (5e-5) — likely undertrained given the branch's residual, near-no-op starting point.
172
+ - Result: persistent artifact (parachute canopy, chain-saw scene) unresolved, but the test wasn't a clean read on whether the context block itself helps, since the branch may not have moved far from its near-identity initialization.
173
+
174
+ ## 10. Next Steps (Queued)
175
+
176
+ - Re-run training with **`--context-lr 3e-4`** as a separate, higher parameter-group LR specifically for the context block, so it actually gets a chance to develop meaningful behavior within the run.
177
+ - Pending that run's results, determine whether the artifact is:
178
+ - Receptive-field-related (context block genuinely helps once properly trained), or
179
+ - Rebalancing/calibration-related (grayscale-confound or lambda tuning), or
180
+ - Caused by something not yet identified.
181
+
182
+ ---
183
+
184
+ ## 11. Tools & Resources Reference
185
+
186
+ | Tool | Purpose |
187
+ |---|---|
188
+ | `hf jobs uv run` | Launch a training/eval/inference job on Hugging Face infrastructure |
189
+ | `hf jobs ps` | List running/recent jobs |
190
+ | `hf jobs logs <job_id>` | Retrieve logs for a job |
191
+ | HF MCP connector (this project) | Exposes `hf_whoami`, `dynamic_space`, `hub_repo_search`, `hub_repo_details`, `hf_fs` (`ls`/`cat`/`stat`/`attach`). **Does not** expose job launching, status, or log retrieval, despite what its documentation/OAuth scope suggests — job management stays CLI-only. |
192
+ | `hf_fs` attach | Retrieves inline images from Hub repos, e.g. `hf://models/User-2468/mini-unet-colorizer/sample.png` |
193
+
194
+ **Model repo:** `User-2468/mini-unet-colorizer`
195
+ **Reliable dataset:** `johnowhitaker/imagenette2-320`
196
+ **Unreliable dataset (avoid):** `frgfm/imagenette`
197
+
198
+ ---
199
+
200
+ *This document reflects the project state as understood from prior conversation history and the current versions of `colorize_train.py`, `colorize_image.py`, and `colorize_eval.py`. It does not include results of the queued `--context-lr 3e-4` run, which had not been executed/reported as of this writing.*
201
+
202
+
203
+ ---
204
+
205
+ ## 12. Verified correction and experiments — 20 September 2026
206
+
207
+ **This section supersedes the earlier root-cause assumptions.** Both main and stable carried stale serialized bin_centers: 232/236 differed from their config. Original training-bin reconstruction exactly matched config (11,943 filtered training images, 236 bins). Shape-only warm-start loading caused target/decode vocabulary disagreement. Repairing the buffer alone reduced reserved 100-image test ab error from 31.27 to 13.90 (55.5%), improving 92/100 images. Learned parameters and raw logits are unchanged; parameter count is 3,968,892.
208
+
209
+ Two controlled 100-update CPU pilots (context LR 5e-5 versus 3e-4; all other conditions equal) completed. Neither beat the repaired baseline on the 50-image validation subset, and both desaturated. Their final checkpoints are retained but rejected for release. The recommended candidate is the mapping-only repair.
210
+
211
+ Inference now supports preserved aspect ratio, original-resolution luminance, EXIF orientation, odd dimensions and explicit revisions. Seven integrity tests passed. Checkpoint loading rejects conflicting bin metadata; new training fixes the vocabulary. New context blocks now actually initialize as identities.
212
+
213
+ The old claim that classification is universally required is too strong. Retain it for this working baseline; other formulations and semantic distillation remain possible research directions. A low-chroma threshold does not prove a source was originally black-and-white. The residual block's mere existence did not make its old initialization an identity.
214
+
215
+ Production readiness is not established: coffee/texture blotches and semantic mistakes remain; diverse external evaluation and broader data are required. Training/test disjointness is verified for the recent seed-0 split, not every older upstream run.
216
+
217
+ Main and stable were not modified: the callable OAuth connection lacked repository-write scope; both direct commits and PR uploads returned 403. A later selected custom Huggingface plugin did not expose tools during the active session. The release candidate and complete experiment bundle are saved with an atomic main upload helper. See `RESEARCH_REPORT.md` for exact evidence, source hashes, split manifests, per-image metrics, limitations and next experiments.
218
+
219
+
220
+ ---
221
+
222
+ ## 13. Spatial research and broader-data training — 20 September 2026
223
+
224
+ Five further runs completed: two 400-update CPU pilots, two matched
225
+ 1,122-update Imagenette GPU arms (classification versus supervised multiscale
226
+ chroma-gradient loss), and a 1,122-update mixed Imagenette/COCO GPU arm.
227
+ The mixed pool contained 17,325 photos. Thirteen spatial decoding variants
228
+ were evaluated. The existing bin mapping was held fixed throughout.
229
+
230
+ The selected release is the mixed-data update-748 checkpoint with single-pass
231
+ luminance-guided filtering, radius 8, epsilon 0.001 and temperature 0.38.
232
+ Two-domain validation selected it after the spatial-loss candidates lost too
233
+ much color on reserved COCO validation photos. On separate Imagenette200 and
234
+ COCO-val100 checks, complete-pipeline chroma error improved 3.54% and 4.81%
235
+ relative to the mapping-only repair. Fine excess-color-edge scores fell
236
+ 85.83% and 84.35%; these are artifact proxies, not human blotch counts. New
237
+ weights improve raw error 1.97% and 3.25% before smoothing.
238
+
239
+ The release retains 3,968,892 parameters. ONNX export includes the full decoder,
240
+ passes dynamic-shape numerical checks, and has a PyTorch-free image wrapper.
241
+ Ten code tests passed. Wrong broad hues remain, including the coffee surface;
242
+ the rocket probe regressed. This is an app-testing candidate, not a claim of
243
+ general production readiness. See RESEARCH_ROUND2.md for selection, tests,
244
+ limits, run IDs and deployment measurements.
245
+
246
+ GPU Jobs are callable through the original Hugging Face connector. The custom
247
+ Huggingface plugin still exposes no tools, and write scope remains absent.
248
+ All GPU weights were recovered through checksum-verified artifact transport;
249
+ best/last weights and experiment records are retained. Main and stable remain
250
+ unchanged. The saved release includes a guarded uploader targeting main.
251
+ The current candidate is release-v2; release/ preserves the earlier repair.
colorize_onnx.py ADDED
@@ -0,0 +1,32 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Run the exported chroma pipeline without importing PyTorch."""
2
+ import argparse
3
+ from pathlib import Path
4
+ import numpy as np
5
+ import onnxruntime as ort
6
+ from PIL import Image,ImageOps
7
+ from skimage.color import rgb2lab,lab2rgb
8
+
9
+ def load_session(path,threads=2):
10
+ opts=ort.SessionOptions();opts.intra_op_num_threads=threads;opts.inter_op_num_threads=1
11
+ return ort.InferenceSession(str(path),sess_options=opts,providers=['CPUExecutionProvider'])
12
+
13
+ def colorize(session,image,size=256):
14
+ if size<8:raise ValueError('size must be at least8')
15
+ image=ImageOps.exif_transpose(image).convert('RGB')
16
+ L=rgb2lab(np.asarray(image,dtype=np.float32)/255)[...,0].astype(np.float32)
17
+ h,w=L.shape;scale=min(size/max(h,w),1)
18
+ target=(max(8,round(w*scale)),max(8,round(h*scale)))
19
+ small=np.asarray(Image.fromarray(L).resize(target,Image.Resampling.BILINEAR),dtype=np.float32)
20
+ x=(small[None,None]/50-1).copy()
21
+ ab=session.run(['chroma'],{'luminance':x})[0][0]
22
+ full=np.stack([np.asarray(Image.fromarray(c).resize((w,h),Image.Resampling.BILINEAR)) for c in ab],axis=-1)
23
+ rgb=np.clip(lab2rgb(np.concatenate([L[...,None],full],axis=-1)),0,1)
24
+ return Image.fromarray(np.rint(rgb*255).astype(np.uint8))
25
+
26
+ if __name__=='__main__':
27
+ p=argparse.ArgumentParser(__doc__);p.add_argument('images',nargs='+');p.add_argument('--model',required=True)
28
+ p.add_argument('--output-dir',default='colorized');p.add_argument('--size',type=int,default=256);p.add_argument('--threads',type=int,default=2)
29
+ a=p.parse_args();session=load_session(a.model,a.threads);out=Path(a.output_dir);out.mkdir(parents=True,exist_ok=True)
30
+ for source in a.images:
31
+ with Image.open(source) as im:result=colorize(session,im,a.size)
32
+ dest=out/(Path(source).stem+'_colorized.png');result.save(dest);print(dest)
colorizer.json ADDED
@@ -0,0 +1,54 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "input": "N,1,H,W Lab L*/50-1; H,W>=8",
3
+ "output": "N,2,H,W Lab chroma a,b",
4
+ "guided_radius": 8,
5
+ "guided_epsilon": 0.001,
6
+ "temperature": 0.38,
7
+ "opset": 17,
8
+ "checks": [
9
+ {
10
+ "shape": [
11
+ 1,
12
+ 1,
13
+ 256,
14
+ 256
15
+ ],
16
+ "max_ab_difference": 8.58306884765625e-06,
17
+ "mean_ab_difference": 1.0279118214384653e-06
18
+ },
19
+ {
20
+ "shape": [
21
+ 1,
22
+ 1,
23
+ 173,
24
+ 241
25
+ ],
26
+ "max_ab_difference": 7.3909759521484375e-06,
27
+ "mean_ab_difference": 1.0783561492644367e-06
28
+ },
29
+ {
30
+ "shape": [
31
+ 2,
32
+ 1,
33
+ 64,
34
+ 80
35
+ ],
36
+ "max_ab_difference": 7.152557373046875e-06,
37
+ "mean_ab_difference": 9.707116532808868e-07
38
+ },
39
+ {
40
+ "shape": [
41
+ 1,
42
+ 1,
43
+ 8,
44
+ 9
45
+ ],
46
+ "max_ab_difference": 1.5497207641601562e-06,
47
+ "mean_ab_difference": 3.6218099808138504e-07
48
+ }
49
+ ],
50
+ "parameters": 3968892,
51
+ "weights_source": "release-v2",
52
+ "runtime": "1.30.0",
53
+ "bytes": 15885757
54
+ }
colorizer.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0ef86749901e66ad53b1e8e1d940330572e2f4c9347ac01e7bc02c7683f8c79a
3
+ size 15885757
export_onnx.py ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Export the colorizer AND luminance-guided chroma decoder as one ONNX graph.
2
+
3
+ Input: N,1,H,W normalized Lab luminance (L*/50-1), minimum8px per side.
4
+ Output: N,2,H,W Lab a,b values. Resize/preserve original L* in the app.
5
+ """
6
+ import argparse,json
7
+ from pathlib import Path
8
+ import numpy as np
9
+ import torch
10
+ from model import load_model
11
+ from spatial import guided_chroma
12
+
13
+ class ChromaPipeline(torch.nn.Module):
14
+ def __init__(self,model,radius=8,temperature=.38):
15
+ super().__init__();self.model=model;self.radius=radius;self.temperature=temperature
16
+ def forward(self,L):
17
+ return guided_chroma(L,self.model.decode(self.model(L),self.temperature),self.radius)
18
+
19
+ def export(a):
20
+ import onnx,onnxruntime as ort
21
+ torch.set_num_threads(2);torch.manual_seed(2026)
22
+ model=ChromaPipeline(load_model(a.model),a.radius).eval()
23
+ out=Path(a.output);out.parent.mkdir(parents=True,exist_ok=True)
24
+ with torch.inference_mode():
25
+ torch.onnx.export(model,torch.zeros(1,1,256,256),str(out),opset_version=17,
26
+ input_names=['luminance'],output_names=['chroma'],dynamo=False,
27
+ dynamic_axes={'luminance':{0:'batch',2:'height',3:'width'},'chroma':{0:'batch',2:'height',3:'width'}})
28
+ onnx.checker.check_model(str(out))
29
+ opts=ort.SessionOptions();opts.intra_op_num_threads=2;opts.inter_op_num_threads=1
30
+ session=ort.InferenceSession(str(out),sess_options=opts,providers=['CPUExecutionProvider'])
31
+ checks=[]
32
+ for shape in [(1,1,256,256),(1,1,173,241),(2,1,64,80),(1,1,8,9)]:
33
+ x=torch.rand(shape)*2-1
34
+ with torch.inference_mode():expected=model(x).numpy()
35
+ actual=session.run(None,{'luminance':x.numpy()})[0]
36
+ error=np.abs(expected-actual)
37
+ assert actual.shape==expected.shape and np.isfinite(actual).all()
38
+ np.testing.assert_allclose(actual,expected,atol=.01,rtol=.001)
39
+ checks.append({'shape':list(shape),'max_ab_difference':float(error.max()),'mean_ab_difference':float(error.mean())})
40
+ result={'input':'N,1,H,W Lab L*/50-1; H,W>=8','output':'N,2,H,W Lab chroma a,b',
41
+ 'guided_radius':a.radius,'guided_epsilon':.001,'temperature':.38,
42
+ 'opset':17,'checks':checks,'parameters':sum(p.numel() for p in model.parameters()),
43
+ 'weights_source':a.model,'runtime':ort.__version__,'bytes':out.stat().st_size}
44
+ out.with_suffix('.json').write_text(json.dumps(result,indent=2));print(json.dumps(result,indent=2))
45
+
46
+ if __name__=='__main__':
47
+ p=argparse.ArgumentParser(__doc__);p.add_argument('--model',required=True);p.add_argument('--output',required=True);p.add_argument('--radius',type=int,default=8)
48
+ export(p.parse_args())
inference.py ADDED
@@ -0,0 +1,63 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Aspect-preserving inference; infer chroma globally and retain original luminance."""
2
+ import argparse
3
+ import math
4
+ from pathlib import Path
5
+ import numpy as np
6
+ from PIL import Image, ImageOps
7
+ import torch
8
+ import torch.nn.functional as F
9
+ from skimage.color import rgb2lab, lab2rgb
10
+ from model import load_model
11
+ from spatial import guided_chroma
12
+
13
+ @torch.inference_mode()
14
+ def colorize(model, image, size=256, temperature=0.38, saturation=1.0, flip_tta=False,
15
+ guided_radius=8, guided_epsilon=.001):
16
+ if size < 8 or not math.isfinite(saturation) or saturation < 0:
17
+ raise ValueError("size must be >=8 and saturation finite and nonnegative")
18
+ image = ImageOps.exif_transpose(image).convert("RGB")
19
+ rgb = np.asarray(image, dtype=np.float32) / 255.0
20
+ luminance = rgb2lab(rgb)[..., 0].astype(np.float32)
21
+ h, w = luminance.shape
22
+ scale = min(size / max(h, w), 1.0)
23
+ target = (max(8, round(h*scale)), max(8, round(w*scale)))
24
+ device = next(model.parameters()).device
25
+ L = torch.from_numpy(luminance)[None, None].to(device) / 50 - 1
26
+ small = F.interpolate(L, size=target, mode="bilinear", align_corners=False, antialias=True)
27
+ logits = model(small)
28
+ if flip_tta:
29
+ logits = (logits + model(small.flip(-1)).flip(-1)) * 0.5
30
+ ab = model.decode(logits, temperature)
31
+ ab = guided_chroma(small, ab, guided_radius, guided_epsilon)
32
+ ab = F.interpolate(ab, size=(h,w), mode="bilinear", align_corners=False)
33
+ ab = ab[0].permute(1,2,0).cpu().numpy() * saturation
34
+ lab = np.concatenate([luminance[...,None], ab], axis=-1)
35
+ result = np.clip(lab2rgb(lab), 0, 1)
36
+ return Image.fromarray(np.rint(result * 255).astype(np.uint8))
37
+
38
+ def main():
39
+ p = argparse.ArgumentParser(__doc__)
40
+ p.add_argument("images", nargs="+")
41
+ p.add_argument("--model", required=True)
42
+ p.add_argument("--revision", default=None)
43
+ p.add_argument("--size", type=int, default=256)
44
+ p.add_argument("--temperature", type=float, default=.38)
45
+ p.add_argument("--saturation", type=float, default=1)
46
+ p.add_argument("--flip-tta", action="store_true")
47
+ p.add_argument("--guided-radius",type=int,default=8,help="Chroma smoothing radius at model resolution;0 disables")
48
+ p.add_argument("--guided-epsilon",type=float,default=.001)
49
+ p.add_argument("--output-dir", default="colorized")
50
+ p.add_argument("--device", default="cuda" if torch.cuda.is_available() else "cpu")
51
+ a = p.parse_args()
52
+ torch.set_num_threads(min(torch.get_num_threads(),4))
53
+ model = load_model(a.model, a.revision, a.device)
54
+ out = Path(a.output_dir); out.mkdir(parents=True, exist_ok=True)
55
+ for file in a.images:
56
+ with Image.open(file) as im:
57
+ result = colorize(model, im, a.size, a.temperature, a.saturation, a.flip_tta,
58
+ a.guided_radius,a.guided_epsilon)
59
+ dest = out / (Path(file).stem + "_colorized.png")
60
+ result.save(dest); print(dest)
61
+
62
+ if __name__ == "__main__":
63
+ main()
model.py ADDED
@@ -0,0 +1,122 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Shared, checkpoint-compatible Mini U-Net. Parameters remain under 4M."""
2
+ import json
3
+ import math
4
+ from pathlib import Path
5
+ import numpy as np
6
+ import torch
7
+ from torch import nn
8
+ import torch.nn.functional as F
9
+ from huggingface_hub import PyTorchModelHubMixin, snapshot_download
10
+ from safetensors.torch import load_file, save_file
11
+ def double_conv(in_ch, out_ch):
12
+ return nn.Sequential(
13
+ nn.Conv2d(in_ch, out_ch, 3, padding=1, bias=False),
14
+ nn.BatchNorm2d(out_ch),
15
+ nn.ReLU(inplace=True),
16
+ nn.Conv2d(out_ch, out_ch, 3, padding=1, bias=False),
17
+ nn.BatchNorm2d(out_ch),
18
+ nn.ReLU(inplace=True),
19
+ )
20
+
21
+
22
+ class DilatedContextBlock(nn.Module):
23
+ def __init__(self, channels, mid_ch=96, dilations=(2, 4, 8)):
24
+ super().__init__()
25
+ self.proj_in = nn.Sequential(
26
+ nn.Conv2d(channels, mid_ch, 1, bias=False),
27
+ nn.BatchNorm2d(mid_ch),
28
+ nn.ReLU(inplace=True),
29
+ )
30
+ layers = []
31
+ for d in dilations:
32
+ layers += [
33
+ nn.Conv2d(mid_ch, mid_ch, 3, padding=d, dilation=d, bias=False),
34
+ nn.BatchNorm2d(mid_ch),
35
+ nn.ReLU(inplace=True),
36
+ ]
37
+ self.dilated = nn.Sequential(*layers)
38
+ self.proj_out = nn.Sequential(
39
+ nn.Conv2d(mid_ch, channels, 1, bias=False),
40
+ nn.BatchNorm2d(channels),
41
+ )
42
+ nn.init.zeros_(self.proj_out[-1].weight)
43
+ self.relu = nn.ReLU(inplace=True)
44
+
45
+ def forward(self, x):
46
+ y = self.proj_in(x)
47
+ y = self.dilated(y)
48
+ y = self.proj_out(y)
49
+ return self.relu(x + y)
50
+
51
+
52
+ class SmallUNetColorizer(
53
+ nn.Module,
54
+ PyTorchModelHubMixin,
55
+ pipeline_tag="image-to-image",
56
+ license="apache-2.0",
57
+ tags=["colorization", "unet", "image-to-image", "classification"],
58
+ ):
59
+ def __init__(self, bin_centers, in_ch: int = 1, base: int = 44,
60
+ context_mid_ch: int = 96, context_dilations=(2, 4, 8)):
61
+ super().__init__()
62
+ self.in_ch, self.base = in_ch, base
63
+ num_bins = len(bin_centers)
64
+ self.num_bins = num_bins
65
+ self.register_buffer("bin_centers", torch.tensor(bin_centers, dtype=torch.float32))
66
+
67
+ self.enc1 = double_conv(in_ch, base)
68
+ self.enc2 = double_conv(base, base * 2)
69
+ self.enc3 = double_conv(base * 2, base * 4)
70
+ self.enc4 = double_conv(base * 4, base * 8)
71
+ self.pool = nn.MaxPool2d(2)
72
+ self.context = DilatedContextBlock(base * 8, mid_ch=context_mid_ch,
73
+ dilations=tuple(context_dilations))
74
+ self.up3 = nn.ConvTranspose2d(base * 8, base * 4, 2, stride=2)
75
+ self.dec3 = double_conv(base * 8, base * 4)
76
+ self.up2 = nn.ConvTranspose2d(base * 4, base * 2, 2, stride=2)
77
+ self.dec2 = double_conv(base * 4, base * 2)
78
+ self.up1 = nn.ConvTranspose2d(base * 2, base, 2, stride=2)
79
+ self.dec1 = double_conv(base * 2, base)
80
+ self.out_conv = nn.Conv2d(base, num_bins, 1)
81
+
82
+ def forward(self, x):
83
+ h, w = x.shape[-2:]
84
+ x = F.pad(x, (0, (-w) % 8, 0, (-h) % 8), mode="replicate")
85
+ e1 = self.enc1(x)
86
+ e2 = self.enc2(self.pool(e1))
87
+ e3 = self.enc3(self.pool(e2))
88
+ e4 = self.context(self.enc4(self.pool(e3)))
89
+ d3 = self.dec3(torch.cat([self.up3(e4), e3], dim=1))
90
+ d2 = self.dec2(torch.cat([self.up2(d3), e2], dim=1))
91
+ d1 = self.dec1(torch.cat([self.up1(d2), e1], dim=1))
92
+ return self.out_conv(d1)[..., :h, :w]
93
+
94
+ def decode(self, logits, temperature: float = 0.38):
95
+ if not math.isfinite(temperature) or temperature <= 0:
96
+ raise ValueError("temperature must be finite and positive")
97
+ probs_t = F.softmax(logits.float() / temperature, dim=1)
98
+ return torch.einsum("bqhw,qc->bchw", probs_t, self.bin_centers)
99
+
100
+
101
+ def load_model(source, revision=None, device="cpu"):
102
+ path = Path(source)
103
+ if not path.is_dir():
104
+ path = Path(snapshot_download(source, revision=revision,
105
+ allow_patterns=["config.json", "model.safetensors"]))
106
+ config = json.loads((path / "config.json").read_text())
107
+ state = load_file(str(path / "model.safetensors"))
108
+ centers = torch.tensor(config["bin_centers"], dtype=torch.float32)
109
+ if not torch.equal(centers, state["bin_centers"]):
110
+ raise ValueError("Checkpoint config and state color bins differ; refusing ambiguous decode")
111
+ model = SmallUNetColorizer(**config)
112
+ model.load_state_dict(state, strict=True)
113
+ model.to(device).eval()
114
+ return model
115
+
116
+ def save_model(model, path):
117
+ path = Path(path); path.mkdir(parents=True, exist_ok=True)
118
+ model.save_pretrained(path)
119
+ # Mixin config can retain constructor bins; use the actual authoritative buffer.
120
+ cfg = json.loads((path / "config.json").read_text())
121
+ cfg["bin_centers"] = model.bin_centers.detach().cpu().tolist()
122
+ (path / "config.json").write_text(json.dumps(cfg, indent=2))
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:16b58fc11dfc3df9c261e4aa1945cf274c32e21aef518817f67ffb0c4a18d288
3
  size 15909320
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0e4c417375684a044860f8af3ac3a2fb44e1a5729254ca33abda758ea71aea6e
3
  size 15909320
reports/round2/coco_manifest.json ADDED
@@ -0,0 +1,1921 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "dataset": "detection-datasets/coco",
3
+ "split": "val",
4
+ "selection": "first100 returned by first-rows; not a random representative sample",
5
+ "images": [
6
+ {
7
+ "index": 0,
8
+ "image_id": 139,
9
+ "path": "000000000139.jpg",
10
+ "sha256": "af55ad0317ee5e22118ca9be7b9989bc6f29058670f7c662a118800f5e481d83",
11
+ "categories": [
12
+ 58,
13
+ 62,
14
+ 62,
15
+ 56,
16
+ 56,
17
+ 56,
18
+ 56,
19
+ 0,
20
+ 0,
21
+ 68,
22
+ 72,
23
+ 73,
24
+ 73,
25
+ 74,
26
+ 75,
27
+ 75,
28
+ 56,
29
+ 75,
30
+ 75,
31
+ 60
32
+ ],
33
+ "size": [
34
+ 640,
35
+ 426
36
+ ]
37
+ },
38
+ {
39
+ "index": 1,
40
+ "image_id": 285,
41
+ "path": "000000000285.jpg",
42
+ "sha256": "257fa1afc7ce320cb502a0f99efbadcc70ff1b37504ab06206e181272c085fb3",
43
+ "categories": [
44
+ 21
45
+ ],
46
+ "size": [
47
+ 586,
48
+ 640
49
+ ]
50
+ },
51
+ {
52
+ "index": 2,
53
+ "image_id": 632,
54
+ "path": "000000000632.jpg",
55
+ "sha256": "9e6cb264ed8e765a0f1c5585b23046e7bbb0a5fe5646cf59885d3bf5231eefb8",
56
+ "categories": [
57
+ 59,
58
+ 58,
59
+ 73,
60
+ 73,
61
+ 73,
62
+ 73,
63
+ 73,
64
+ 56,
65
+ 58,
66
+ 73,
67
+ 73,
68
+ 73,
69
+ 73,
70
+ 73,
71
+ 73,
72
+ 73,
73
+ 73,
74
+ 73
75
+ ],
76
+ "size": [
77
+ 640,
78
+ 483
79
+ ]
80
+ },
81
+ {
82
+ "index": 3,
83
+ "image_id": 724,
84
+ "path": "000000000724.jpg",
85
+ "sha256": "58d4ecf6b28bafd8d84fe27545f2d2e4030bf32c959d4b53520b7ffa61b42fcd",
86
+ "categories": [
87
+ 11,
88
+ 7,
89
+ 2,
90
+ 11
91
+ ],
92
+ "size": [
93
+ 375,
94
+ 500
95
+ ]
96
+ },
97
+ {
98
+ "index": 4,
99
+ "image_id": 776,
100
+ "path": "000000000776.jpg",
101
+ "sha256": "f89815e27c0b08b51316776c1c53e3871644086f979fcf138b99a2a45cafd6ac",
102
+ "categories": [
103
+ 77,
104
+ 77,
105
+ 77,
106
+ 59
107
+ ],
108
+ "size": [
109
+ 428,
110
+ 640
111
+ ]
112
+ },
113
+ {
114
+ "index": 5,
115
+ "image_id": 785,
116
+ "path": "000000000785.jpg",
117
+ "sha256": "70df38aace8fb6bb296d2c5564b4a6ea7b34488c30e353fa1fb69cf46d753987",
118
+ "categories": [
119
+ 0,
120
+ 30
121
+ ],
122
+ "size": [
123
+ 640,
124
+ 425
125
+ ]
126
+ },
127
+ {
128
+ "index": 6,
129
+ "image_id": 802,
130
+ "path": "000000000802.jpg",
131
+ "sha256": "d089b07aff5346bf76257b1228f68bd30e5e03535b11fc3dc7dc0ff86a928ace",
132
+ "categories": [
133
+ 72,
134
+ 69
135
+ ],
136
+ "size": [
137
+ 424,
138
+ 640
139
+ ]
140
+ },
141
+ {
142
+ "index": 7,
143
+ "image_id": 872,
144
+ "path": "000000000872.jpg",
145
+ "sha256": "af1c420a74cb563fac48ff9a7d3e42298e6f4c2e534dea16e70a96941c440107",
146
+ "categories": [
147
+ 32,
148
+ 0,
149
+ 0,
150
+ 35
151
+ ],
152
+ "size": [
153
+ 621,
154
+ 640
155
+ ]
156
+ },
157
+ {
158
+ "index": 8,
159
+ "image_id": 885,
160
+ "path": "000000000885.jpg",
161
+ "sha256": "a90b60cf465d373b13fae02a500aec9ce6e4a7e13194e59d594848aa4b3eb661",
162
+ "categories": [
163
+ 0,
164
+ 0,
165
+ 0,
166
+ 0,
167
+ 0,
168
+ 38,
169
+ 0,
170
+ 0,
171
+ 0
172
+ ],
173
+ "size": [
174
+ 640,
175
+ 427
176
+ ]
177
+ },
178
+ {
179
+ "index": 9,
180
+ "image_id": 1000,
181
+ "path": "000000001000.jpg",
182
+ "sha256": "52db327549029145b7c9d4bf6fd4366750d227f261fedbdfe30caea7485d2131",
183
+ "categories": [
184
+ 38,
185
+ 26,
186
+ 26,
187
+ 0,
188
+ 0,
189
+ 0,
190
+ 0,
191
+ 0,
192
+ 0,
193
+ 0,
194
+ 0,
195
+ 0,
196
+ 0,
197
+ 24,
198
+ 24,
199
+ 0,
200
+ 0
201
+ ],
202
+ "size": [
203
+ 640,
204
+ 480
205
+ ]
206
+ },
207
+ {
208
+ "index": 10,
209
+ "image_id": 1268,
210
+ "path": "000000001268.jpg",
211
+ "sha256": "5bc27557eecc9e6da0b77cfe4db1f8f60515044454f27833553242e2171b17f1",
212
+ "categories": [
213
+ 14,
214
+ 8,
215
+ 8,
216
+ 0,
217
+ 0,
218
+ 0,
219
+ 0,
220
+ 67,
221
+ 24,
222
+ 26,
223
+ 8
224
+ ],
225
+ "size": [
226
+ 640,
227
+ 427
228
+ ]
229
+ },
230
+ {
231
+ "index": 11,
232
+ "image_id": 1296,
233
+ "path": "000000001296.jpg",
234
+ "sha256": "3ad6b3e0ac9d3a2c23dd68d4fa1022a9925c65b1f8b7ae63fe13cf81c25e134e",
235
+ "categories": [
236
+ 67,
237
+ 74,
238
+ 0,
239
+ 0
240
+ ],
241
+ "size": [
242
+ 427,
243
+ 640
244
+ ]
245
+ },
246
+ {
247
+ "index": 12,
248
+ "image_id": 1353,
249
+ "path": "000000001353.jpg",
250
+ "sha256": "f7ad2e785db744e89b1c515074c3dba6d025352370bc5cebc50c72947f7fd5f1",
251
+ "categories": [
252
+ 6,
253
+ 0,
254
+ 0,
255
+ 0,
256
+ 0,
257
+ 0,
258
+ 0
259
+ ],
260
+ "size": [
261
+ 375,
262
+ 500
263
+ ]
264
+ },
265
+ {
266
+ "index": 13,
267
+ "image_id": 1425,
268
+ "path": "000000001425.jpg",
269
+ "sha256": "2142fa8eecb43eac446ff9e7e2fea2411396daba7f1fb358eeed6ef056193201",
270
+ "categories": [
271
+ 48,
272
+ 45
273
+ ],
274
+ "size": [
275
+ 640,
276
+ 512
277
+ ]
278
+ },
279
+ {
280
+ "index": 14,
281
+ "image_id": 1490,
282
+ "path": "000000001490.jpg",
283
+ "sha256": "390462cb0c6adf324280cee48fa1cc4b2de6251e6f3cb85c6954216928edd42f",
284
+ "categories": [
285
+ 0,
286
+ 37
287
+ ],
288
+ "size": [
289
+ 640,
290
+ 315
291
+ ]
292
+ },
293
+ {
294
+ "index": 15,
295
+ "image_id": 1503,
296
+ "path": "000000001503.jpg",
297
+ "sha256": "6387d2f23e7049d7baf30fcc374bb866174e5ef90b5e41ed09ad587f6d22c723",
298
+ "categories": [
299
+ 63,
300
+ 64,
301
+ 66,
302
+ 62,
303
+ 64
304
+ ],
305
+ "size": [
306
+ 320,
307
+ 240
308
+ ]
309
+ },
310
+ {
311
+ "index": 16,
312
+ "image_id": 1532,
313
+ "path": "000000001532.jpg",
314
+ "sha256": "3d673d6588ce195e9bb9b5da570abf1cd386589b0a5e6314cfe5aa7efd617c59",
315
+ "categories": [
316
+ 2,
317
+ 2,
318
+ 2,
319
+ 2,
320
+ 7,
321
+ 2,
322
+ 2,
323
+ 2
324
+ ],
325
+ "size": [
326
+ 640,
327
+ 480
328
+ ]
329
+ },
330
+ {
331
+ "index": 17,
332
+ "image_id": 1584,
333
+ "path": "000000001584.jpg",
334
+ "sha256": "c4a7561884c6872a29da0564f2dee4a885fd0da6ba902a7d46e5052f677f5c70",
335
+ "categories": [
336
+ 5,
337
+ 0,
338
+ 0,
339
+ 0,
340
+ 0,
341
+ 0,
342
+ 0,
343
+ 5,
344
+ 0,
345
+ 0,
346
+ 0,
347
+ 0,
348
+ 0,
349
+ 5
350
+ ],
351
+ "size": [
352
+ 612,
353
+ 612
354
+ ]
355
+ },
356
+ {
357
+ "index": 18,
358
+ "image_id": 1675,
359
+ "path": "000000001675.jpg",
360
+ "sha256": "969bfc8e03ff8404287a541a49f7f92b0e010082a9af443893834de78dcf91a5",
361
+ "categories": [
362
+ 15,
363
+ 66
364
+ ],
365
+ "size": [
366
+ 640,
367
+ 480
368
+ ]
369
+ },
370
+ {
371
+ "index": 19,
372
+ "image_id": 1761,
373
+ "path": "000000001761.jpg",
374
+ "sha256": "8b07a2fff9a4207bf77da6ea9dd01985363bd9c66a6c519b70aa765ae7e4cea5",
375
+ "categories": [
376
+ 4,
377
+ 4,
378
+ 0,
379
+ 0,
380
+ 0,
381
+ 0,
382
+ 0
383
+ ],
384
+ "size": [
385
+ 427,
386
+ 640
387
+ ]
388
+ },
389
+ {
390
+ "index": 20,
391
+ "image_id": 1818,
392
+ "path": "000000001818.jpg",
393
+ "sha256": "99f2a4fb51a38407d0adaa6807fffe0ff4a089fb5362cf5e4496b2d0734a3700",
394
+ "categories": [
395
+ 22,
396
+ 22
397
+ ],
398
+ "size": [
399
+ 640,
400
+ 425
401
+ ]
402
+ },
403
+ {
404
+ "index": 21,
405
+ "image_id": 1993,
406
+ "path": "000000001993.jpg",
407
+ "sha256": "83fe84fc8bd4679bb1f71343b9747b67252ce04cc375a1c4ec477d9150d27475",
408
+ "categories": [
409
+ 59,
410
+ 56,
411
+ 56,
412
+ 60
413
+ ],
414
+ "size": [
415
+ 640,
416
+ 419
417
+ ]
418
+ },
419
+ {
420
+ "index": 22,
421
+ "image_id": 2006,
422
+ "path": "000000002006.jpg",
423
+ "sha256": "0c83931a1194297853ac8058f3e80cb16f59ede4838a05dd37ca90db39733e1e",
424
+ "categories": [
425
+ 5,
426
+ 0,
427
+ 0,
428
+ 27,
429
+ 27,
430
+ 0,
431
+ 9,
432
+ 9
433
+ ],
434
+ "size": [
435
+ 640,
436
+ 480
437
+ ]
438
+ },
439
+ {
440
+ "index": 23,
441
+ "image_id": 2149,
442
+ "path": "000000002149.jpg",
443
+ "sha256": "18feb34fd5b7241fd2fb97918c503caaa6992e9b82b397c8d213ffd6b7246b99",
444
+ "categories": [
445
+ 45,
446
+ 47
447
+ ],
448
+ "size": [
449
+ 640,
450
+ 427
451
+ ]
452
+ },
453
+ {
454
+ "index": 24,
455
+ "image_id": 2153,
456
+ "path": "000000002153.jpg",
457
+ "sha256": "97a1a88f9c56fd18eb3bc98ec51ff5437c4f2aa48508ded248bf7bc70f9d04ff",
458
+ "categories": [
459
+ 0,
460
+ 0,
461
+ 34,
462
+ 0,
463
+ 0
464
+ ],
465
+ "size": [
466
+ 640,
467
+ 480
468
+ ]
469
+ },
470
+ {
471
+ "index": 25,
472
+ "image_id": 2157,
473
+ "path": "000000002157.jpg",
474
+ "sha256": "250d45d50a97d995bf8afb564b0d3c807fb5ac2c8fd92927c990cbdae4f52198",
475
+ "categories": [
476
+ 43,
477
+ 43,
478
+ 55,
479
+ 40,
480
+ 40,
481
+ 40,
482
+ 40,
483
+ 40,
484
+ 40,
485
+ 40,
486
+ 40,
487
+ 40,
488
+ 41,
489
+ 41,
490
+ 60,
491
+ 43
492
+ ],
493
+ "size": [
494
+ 640,
495
+ 427
496
+ ]
497
+ },
498
+ {
499
+ "index": 26,
500
+ "image_id": 2261,
501
+ "path": "000000002261.jpg",
502
+ "sha256": "070019cc062447bebe627443c764b20f9af4ae0cc4788ff16b32776717274bb7",
503
+ "categories": [
504
+ 0,
505
+ 37
506
+ ],
507
+ "size": [
508
+ 640,
509
+ 427
510
+ ]
511
+ },
512
+ {
513
+ "index": 27,
514
+ "image_id": 2299,
515
+ "path": "000000002299.jpg",
516
+ "sha256": "3fe991ca6fcc6fda312dfc77230b5f4d0330aa23dc6bda1b42d676ca785c0c80",
517
+ "categories": [
518
+ 27,
519
+ 27,
520
+ 27,
521
+ 27,
522
+ 0,
523
+ 0,
524
+ 0,
525
+ 0,
526
+ 0,
527
+ 0,
528
+ 0,
529
+ 0,
530
+ 0,
531
+ 0,
532
+ 0,
533
+ 0,
534
+ 27,
535
+ 27,
536
+ 27,
537
+ 27,
538
+ 27,
539
+ 0,
540
+ 0
541
+ ],
542
+ "size": [
543
+ 500,
544
+ 302
545
+ ]
546
+ },
547
+ {
548
+ "index": 28,
549
+ "image_id": 2431,
550
+ "path": "000000002431.jpg",
551
+ "sha256": "6299a9bf9f37cd4daf5aa2eedeba0fc9c89a44df804499e1f27844c2179d269a",
552
+ "categories": [
553
+ 40,
554
+ 41,
555
+ 41,
556
+ 43,
557
+ 44,
558
+ 44,
559
+ 60,
560
+ 0,
561
+ 0
562
+ ],
563
+ "size": [
564
+ 457,
565
+ 640
566
+ ]
567
+ },
568
+ {
569
+ "index": 29,
570
+ "image_id": 2473,
571
+ "path": "000000002473.jpg",
572
+ "sha256": "4dff42d731df540ccc1c5861710f322f0bd17118d054a19771c7ba18f8ac3cfc",
573
+ "categories": [
574
+ 0,
575
+ 0,
576
+ 0,
577
+ 0,
578
+ 30
579
+ ],
580
+ "size": [
581
+ 640,
582
+ 427
583
+ ]
584
+ },
585
+ {
586
+ "index": 30,
587
+ "image_id": 2532,
588
+ "path": "000000002532.jpg",
589
+ "sha256": "b1f09e4286caa56b97c007ac74a5689bcfe1edbed9c31b3dfe9ca82f561f1511",
590
+ "categories": [
591
+ 0,
592
+ 30
593
+ ],
594
+ "size": [
595
+ 480,
596
+ 640
597
+ ]
598
+ },
599
+ {
600
+ "index": 31,
601
+ "image_id": 2587,
602
+ "path": "000000002587.jpg",
603
+ "sha256": "e6222faf896e4b78f24135a5a939abaaa838d26a7657d9182405fa90a0e64bde",
604
+ "categories": [
605
+ 46,
606
+ 54
607
+ ],
608
+ "size": [
609
+ 500,
610
+ 375
611
+ ]
612
+ },
613
+ {
614
+ "index": 32,
615
+ "image_id": 2592,
616
+ "path": "000000002592.jpg",
617
+ "sha256": "66c6584eb97aabcce3ad1297aaf9352a294f7a90a4ef5da3cd488501165a6471",
618
+ "categories": [
619
+ 60,
620
+ 41,
621
+ 43
622
+ ],
623
+ "size": [
624
+ 640,
625
+ 366
626
+ ]
627
+ },
628
+ {
629
+ "index": 33,
630
+ "image_id": 2685,
631
+ "path": "000000002685.jpg",
632
+ "sha256": "4ae6fa81410eecf977485edd76c05b85504f02a26689e446077b7a115d4ce2d2",
633
+ "categories": [
634
+ 39,
635
+ 39,
636
+ 39,
637
+ 39,
638
+ 39,
639
+ 39,
640
+ 39,
641
+ 0,
642
+ 0,
643
+ 0,
644
+ 0,
645
+ 39,
646
+ 40,
647
+ 41,
648
+ 0,
649
+ 26,
650
+ 40,
651
+ 41,
652
+ 0
653
+ ],
654
+ "size": [
655
+ 640,
656
+ 555
657
+ ]
658
+ },
659
+ {
660
+ "index": 34,
661
+ "image_id": 2923,
662
+ "path": "000000002923.jpg",
663
+ "sha256": "cc86aff672222b01de96336c0952695d5b1ab442f2cede77bd3a305674b68117",
664
+ "categories": [
665
+ 14,
666
+ 14,
667
+ 14,
668
+ 8,
669
+ 8,
670
+ 8,
671
+ 8
672
+ ],
673
+ "size": [
674
+ 500,
675
+ 375
676
+ ]
677
+ },
678
+ {
679
+ "index": 35,
680
+ "image_id": 3156,
681
+ "path": "000000003156.jpg",
682
+ "sha256": "3f4967b7da9d1362e8a23487a2aebb0fd79ad62b447047c90daf23a7c88e281a",
683
+ "categories": [
684
+ 0,
685
+ 71,
686
+ 61
687
+ ],
688
+ "size": [
689
+ 443,
690
+ 640
691
+ ]
692
+ },
693
+ {
694
+ "index": 36,
695
+ "image_id": 3255,
696
+ "path": "000000003255.jpg",
697
+ "sha256": "b620ca00c8ffa7d38cc5d54f3efdb0c36e2a94801719f1173463070f395b0da9",
698
+ "categories": [
699
+ 0,
700
+ 0,
701
+ 0,
702
+ 0,
703
+ 30,
704
+ 24,
705
+ 24,
706
+ 0,
707
+ 24,
708
+ 0
709
+ ],
710
+ "size": [
711
+ 640,
712
+ 363
713
+ ]
714
+ },
715
+ {
716
+ "index": 37,
717
+ "image_id": 3501,
718
+ "path": "000000003501.jpg",
719
+ "sha256": "872beddd8e45dd2fbceb10ff6463a73b7be3c1b2dcefaa827386ffedd30eebdf",
720
+ "categories": [
721
+ 50,
722
+ 50,
723
+ 45
724
+ ],
725
+ "size": [
726
+ 612,
727
+ 612
728
+ ]
729
+ },
730
+ {
731
+ "index": 38,
732
+ "image_id": 3553,
733
+ "path": "000000003553.jpg",
734
+ "sha256": "f0e8602875c32ceb441ce5bbd1fc1e9d4ee9116c3f97113a8167a86a0b6f2c6e",
735
+ "categories": [
736
+ 0,
737
+ 36
738
+ ],
739
+ "size": [
740
+ 640,
741
+ 427
742
+ ]
743
+ },
744
+ {
745
+ "index": 39,
746
+ "image_id": 3661,
747
+ "path": "000000003661.jpg",
748
+ "sha256": "91580743377d827fb505726da7ae3c4311a7a8826979d64978e308cc832b0b1f",
749
+ "categories": [
750
+ 46,
751
+ 66,
752
+ 41
753
+ ],
754
+ "size": [
755
+ 640,
756
+ 384
757
+ ]
758
+ },
759
+ {
760
+ "index": 40,
761
+ "image_id": 3845,
762
+ "path": "000000003845.jpg",
763
+ "sha256": "ba8e79b63ee006c3242543bfdcb6e7a34b1c5640ba31d8283a9cb86d98e05cc3",
764
+ "categories": [
765
+ 60,
766
+ 41,
767
+ 42,
768
+ 50,
769
+ 50,
770
+ 50,
771
+ 51,
772
+ 51,
773
+ 51,
774
+ 51,
775
+ 44,
776
+ 51
777
+ ],
778
+ "size": [
779
+ 500,
780
+ 375
781
+ ]
782
+ },
783
+ {
784
+ "index": 41,
785
+ "image_id": 3934,
786
+ "path": "000000003934.jpg",
787
+ "sha256": "f7f77f443b29ac83e363901a4e1259e042196af39b66e1e943c0e89156926b07",
788
+ "categories": [
789
+ 57,
790
+ 0,
791
+ 0,
792
+ 0,
793
+ 41,
794
+ 65,
795
+ 65,
796
+ 0,
797
+ 0,
798
+ 40,
799
+ 40,
800
+ 40,
801
+ 0,
802
+ 0
803
+ ],
804
+ "size": [
805
+ 375,
806
+ 500
807
+ ]
808
+ },
809
+ {
810
+ "index": 42,
811
+ "image_id": 4134,
812
+ "path": "000000004134.jpg",
813
+ "sha256": "8e8bea5c2f74fbc3ffb1654f8ad0172d76b5d3b337417b09228d1a93d7fc03dd",
814
+ "categories": [
815
+ 27,
816
+ 0,
817
+ 0,
818
+ 0,
819
+ 0,
820
+ 0,
821
+ 0,
822
+ 0,
823
+ 0,
824
+ 0,
825
+ 40,
826
+ 40,
827
+ 40,
828
+ 60,
829
+ 0,
830
+ 27,
831
+ 27,
832
+ 40,
833
+ 56,
834
+ 56,
835
+ 0,
836
+ 0,
837
+ 56,
838
+ 60,
839
+ 0,
840
+ 60
841
+ ],
842
+ "size": [
843
+ 640,
844
+ 425
845
+ ]
846
+ },
847
+ {
848
+ "index": 43,
849
+ "image_id": 4395,
850
+ "path": "000000004395.jpg",
851
+ "sha256": "d0123d4b85b750596124de86fdc37132c5c7f5a84eb6cf4d178f60cacc8a4daf",
852
+ "categories": [
853
+ 27,
854
+ 0
855
+ ],
856
+ "size": [
857
+ 425,
858
+ 640
859
+ ]
860
+ },
861
+ {
862
+ "index": 44,
863
+ "image_id": 4495,
864
+ "path": "000000004495.jpg",
865
+ "sha256": "8bcab9c4070640d8ac4f99809cab84c60e7baf422bf17fc136d5416d2763547e",
866
+ "categories": [
867
+ 62,
868
+ 57,
869
+ 56
870
+ ],
871
+ "size": [
872
+ 500,
873
+ 375
874
+ ]
875
+ },
876
+ {
877
+ "index": 45,
878
+ "image_id": 4765,
879
+ "path": "000000004765.jpg",
880
+ "sha256": "4461e72ed180471bf0330bb8f57a829a5215d698be7f876f733b165c9d2ae0b1",
881
+ "categories": [
882
+ 0,
883
+ 37
884
+ ],
885
+ "size": [
886
+ 612,
887
+ 612
888
+ ]
889
+ },
890
+ {
891
+ "index": 46,
892
+ "image_id": 4795,
893
+ "path": "000000004795.jpg",
894
+ "sha256": "d232b0316193e98ea6da541b21059e4466a37868d605057fb4aca55ff6ef2f19",
895
+ "categories": [
896
+ 15,
897
+ 63,
898
+ 62
899
+ ],
900
+ "size": [
901
+ 640,
902
+ 480
903
+ ]
904
+ },
905
+ {
906
+ "index": 47,
907
+ "image_id": 5001,
908
+ "path": "000000005001.jpg",
909
+ "sha256": "1b9e8598ca2fc43b263cd848e72b9e1b2ae45cf28d36b30715f3b3072e72bede",
910
+ "categories": [
911
+ 76,
912
+ 0,
913
+ 0,
914
+ 0,
915
+ 0,
916
+ 0,
917
+ 0,
918
+ 0,
919
+ 0,
920
+ 0,
921
+ 0,
922
+ 0,
923
+ 0,
924
+ 26,
925
+ 1,
926
+ 0,
927
+ 76,
928
+ 0
929
+ ],
930
+ "size": [
931
+ 640,
932
+ 480
933
+ ]
934
+ },
935
+ {
936
+ "index": 48,
937
+ "image_id": 5037,
938
+ "path": "000000005037.jpg",
939
+ "sha256": "16b03f54c20f81798d2a281b2dc568b5ddf28e8831dfa6dbbe3335d170a82fb2",
940
+ "categories": [
941
+ 5,
942
+ 2,
943
+ 0,
944
+ 0,
945
+ 0,
946
+ 0,
947
+ 0,
948
+ 0,
949
+ 0,
950
+ 2
951
+ ],
952
+ "size": [
953
+ 640,
954
+ 425
955
+ ]
956
+ },
957
+ {
958
+ "index": 49,
959
+ "image_id": 5060,
960
+ "path": "000000005060.jpg",
961
+ "sha256": "1e7981ba6f351983fbc75e37d63724660921be559d48ea07189935c5dec038f7",
962
+ "categories": [
963
+ 67,
964
+ 0
965
+ ],
966
+ "size": [
967
+ 480,
968
+ 640
969
+ ]
970
+ },
971
+ {
972
+ "index": 50,
973
+ "image_id": 5193,
974
+ "path": "000000005193.jpg",
975
+ "sha256": "d5efc50e04c9de9046679fcd40a8a57b8006160874fe78359ea8c5b6d0b78332",
976
+ "categories": [
977
+ 0,
978
+ 0,
979
+ 0,
980
+ 0,
981
+ 0,
982
+ 37,
983
+ 37,
984
+ 39,
985
+ 0
986
+ ],
987
+ "size": [
988
+ 640,
989
+ 425
990
+ ]
991
+ },
992
+ {
993
+ "index": 51,
994
+ "image_id": 5477,
995
+ "path": "000000005477.jpg",
996
+ "sha256": "353d735f9ad7d4be09382c8d8e65ca252a98fefd755406331b50fd077a31b9c3",
997
+ "categories": [
998
+ 4,
999
+ 4
1000
+ ],
1001
+ "size": [
1002
+ 640,
1003
+ 349
1004
+ ]
1005
+ },
1006
+ {
1007
+ "index": 52,
1008
+ "image_id": 5503,
1009
+ "path": "000000005503.jpg",
1010
+ "sha256": "23cc0f458dc6488f9fdb8ea6c46a5ac1e50cdc1de4e82f4146cfb8904363c816",
1011
+ "categories": [
1012
+ 61
1013
+ ],
1014
+ "size": [
1015
+ 333,
1016
+ 500
1017
+ ]
1018
+ },
1019
+ {
1020
+ "index": 53,
1021
+ "image_id": 5529,
1022
+ "path": "000000005529.jpg",
1023
+ "sha256": "3df22cce0c2f1eece4aea3c5abd0add7081ecd7deeaed7df2377ce9d3321811c",
1024
+ "categories": [
1025
+ 0,
1026
+ 30
1027
+ ],
1028
+ "size": [
1029
+ 444,
1030
+ 640
1031
+ ]
1032
+ },
1033
+ {
1034
+ "index": 54,
1035
+ "image_id": 5586,
1036
+ "path": "000000005586.jpg",
1037
+ "sha256": "f85564715c5ecdde28bb0340458aae4f86f1be9689f4c14309474abb3a10b9c8",
1038
+ "categories": [
1039
+ 0,
1040
+ 0,
1041
+ 38,
1042
+ 0,
1043
+ 0,
1044
+ 0,
1045
+ 0,
1046
+ 0,
1047
+ 0,
1048
+ 0,
1049
+ 0,
1050
+ 0,
1051
+ 0,
1052
+ 0,
1053
+ 0
1054
+ ],
1055
+ "size": [
1056
+ 320,
1057
+ 240
1058
+ ]
1059
+ },
1060
+ {
1061
+ "index": 55,
1062
+ "image_id": 5600,
1063
+ "path": "000000005600.jpg",
1064
+ "sha256": "a9686c450bf48766888a2fe8be3c00cd6cceb4c373c25d17aadc5e1cf635a89d",
1065
+ "categories": [
1066
+ 44,
1067
+ 45,
1068
+ 45
1069
+ ],
1070
+ "size": [
1071
+ 640,
1072
+ 361
1073
+ ]
1074
+ },
1075
+ {
1076
+ "index": 56,
1077
+ "image_id": 5992,
1078
+ "path": "000000005992.jpg",
1079
+ "sha256": "d372f589b9bac7b6f8f9cfd5ce02e571d878fb440422b6630f3e720dc35fee20",
1080
+ "categories": [
1081
+ 18,
1082
+ 18,
1083
+ 18,
1084
+ 18,
1085
+ 18
1086
+ ],
1087
+ "size": [
1088
+ 640,
1089
+ 428
1090
+ ]
1091
+ },
1092
+ {
1093
+ "index": 57,
1094
+ "image_id": 6012,
1095
+ "path": "000000006012.jpg",
1096
+ "sha256": "1e3b6551efe47e0b6042eaef93144556945d9c09537465c480bfeabbf0925233",
1097
+ "categories": [
1098
+ 46,
1099
+ 46
1100
+ ],
1101
+ "size": [
1102
+ 640,
1103
+ 531
1104
+ ]
1105
+ },
1106
+ {
1107
+ "index": 58,
1108
+ "image_id": 6040,
1109
+ "path": "000000006040.jpg",
1110
+ "sha256": "0f54fdcbf99d794e363f452a2a414abba9dfe29d388320caa9a34dcfa93134f2",
1111
+ "categories": [
1112
+ 6,
1113
+ 0,
1114
+ 0,
1115
+ 0,
1116
+ 2,
1117
+ 0,
1118
+ 0,
1119
+ 0,
1120
+ 0,
1121
+ 0,
1122
+ 0,
1123
+ 7
1124
+ ],
1125
+ "size": [
1126
+ 640,
1127
+ 351
1128
+ ]
1129
+ },
1130
+ {
1131
+ "index": 59,
1132
+ "image_id": 6213,
1133
+ "path": "000000006213.jpg",
1134
+ "sha256": "3f8b5aa95567370b60167842fa18b5b2cb81cd9315aa87402f9410d10b8929b3",
1135
+ "categories": [
1136
+ 71,
1137
+ 71
1138
+ ],
1139
+ "size": [
1140
+ 640,
1141
+ 425
1142
+ ]
1143
+ },
1144
+ {
1145
+ "index": 60,
1146
+ "image_id": 6460,
1147
+ "path": "000000006460.jpg",
1148
+ "sha256": "f2d01a1a2c0494390649feaa5ac37c616dfcff34a1210421e8c3e012e6f415f7",
1149
+ "categories": [
1150
+ 37,
1151
+ 0
1152
+ ],
1153
+ "size": [
1154
+ 640,
1155
+ 427
1156
+ ]
1157
+ },
1158
+ {
1159
+ "index": 61,
1160
+ "image_id": 6471,
1161
+ "path": "000000006471.jpg",
1162
+ "sha256": "854c60de96d0adf7d2864a3026195dad58f30cde7d694f743640b65da8ed2697",
1163
+ "categories": [
1164
+ 0,
1165
+ 0,
1166
+ 0,
1167
+ 0,
1168
+ 0,
1169
+ 0,
1170
+ 34,
1171
+ 0,
1172
+ 13,
1173
+ 13,
1174
+ 35,
1175
+ 39,
1176
+ 0,
1177
+ 0,
1178
+ 39,
1179
+ 0
1180
+ ],
1181
+ "size": [
1182
+ 500,
1183
+ 333
1184
+ ]
1185
+ },
1186
+ {
1187
+ "index": 62,
1188
+ "image_id": 6614,
1189
+ "path": "000000006614.jpg",
1190
+ "sha256": "90003c6f2d34589d375819427576891341cf395d1d24b63ab85ffc959a037af9",
1191
+ "categories": [
1192
+ 47,
1193
+ 49
1194
+ ],
1195
+ "size": [
1196
+ 500,
1197
+ 396
1198
+ ]
1199
+ },
1200
+ {
1201
+ "index": 63,
1202
+ "image_id": 6723,
1203
+ "path": "000000006723.jpg",
1204
+ "sha256": "a1106507bc1bc4756e525efdc24a4c5361ecf6116ca0fea5c48cbf44393cba05",
1205
+ "categories": [
1206
+ 2,
1207
+ 2,
1208
+ 2,
1209
+ 5,
1210
+ 7,
1211
+ 7
1212
+ ],
1213
+ "size": [
1214
+ 640,
1215
+ 361
1216
+ ]
1217
+ },
1218
+ {
1219
+ "index": 64,
1220
+ "image_id": 6763,
1221
+ "path": "000000006763.jpg",
1222
+ "sha256": "80769a0536ee73cd3abd3b267f95f2df8ac56d7d2b56713afb5ce65ce0b0be48",
1223
+ "categories": [
1224
+ 62,
1225
+ 0,
1226
+ 0,
1227
+ 27,
1228
+ 60,
1229
+ 67,
1230
+ 0
1231
+ ],
1232
+ "size": [
1233
+ 375,
1234
+ 500
1235
+ ]
1236
+ },
1237
+ {
1238
+ "index": 65,
1239
+ "image_id": 6771,
1240
+ "path": "000000006771.jpg",
1241
+ "sha256": "1d47e3ca98f3f76f6206eadabc4c984400f737fcf55b59ad731dccb1df42b454",
1242
+ "categories": [
1243
+ 67,
1244
+ 0,
1245
+ 0,
1246
+ 0,
1247
+ 0,
1248
+ 0,
1249
+ 0,
1250
+ 0,
1251
+ 0
1252
+ ],
1253
+ "size": [
1254
+ 640,
1255
+ 427
1256
+ ]
1257
+ },
1258
+ {
1259
+ "index": 66,
1260
+ "image_id": 6818,
1261
+ "path": "000000006818.jpg",
1262
+ "sha256": "fb502df748f83d4d3c61d9c20cb9acaec4f3c967f3d451861fbdfc894956fe15",
1263
+ "categories": [
1264
+ 61
1265
+ ],
1266
+ "size": [
1267
+ 427,
1268
+ 640
1269
+ ]
1270
+ },
1271
+ {
1272
+ "index": 67,
1273
+ "image_id": 6894,
1274
+ "path": "000000006894.jpg",
1275
+ "sha256": "c648b3be74ed897658bf6925377adf94b536b33b249fb7caa72cf6243d991e40",
1276
+ "categories": [
1277
+ 0,
1278
+ 20
1279
+ ],
1280
+ "size": [
1281
+ 640,
1282
+ 480
1283
+ ]
1284
+ },
1285
+ {
1286
+ "index": 68,
1287
+ "image_id": 6954,
1288
+ "path": "000000006954.jpg",
1289
+ "sha256": "d0deb0fd648f581c6808e2a5e64a0bb38159704389d605382c4ecfc5a31feabd",
1290
+ "categories": [
1291
+ 0,
1292
+ 0,
1293
+ 0,
1294
+ 0,
1295
+ 0,
1296
+ 29,
1297
+ 29
1298
+ ],
1299
+ "size": [
1300
+ 640,
1301
+ 480
1302
+ ]
1303
+ },
1304
+ {
1305
+ "index": 69,
1306
+ "image_id": 7088,
1307
+ "path": "000000007088.jpg",
1308
+ "sha256": "df2ffa1e67f0245d6310b4d16966d21fad9e73fbadeb0ce525837448e5f2b5a6",
1309
+ "categories": [
1310
+ 25,
1311
+ 7,
1312
+ 0,
1313
+ 2
1314
+ ],
1315
+ "size": [
1316
+ 478,
1317
+ 640
1318
+ ]
1319
+ },
1320
+ {
1321
+ "index": 70,
1322
+ "image_id": 7108,
1323
+ "path": "000000007108.jpg",
1324
+ "sha256": "44c12a5f0306cef880c577b03afc68d53ab5117a496c90b5019efe8660ee4dcc",
1325
+ "categories": [
1326
+ 20,
1327
+ 20,
1328
+ 20,
1329
+ 20,
1330
+ 20
1331
+ ],
1332
+ "size": [
1333
+ 640,
1334
+ 426
1335
+ ]
1336
+ },
1337
+ {
1338
+ "index": 71,
1339
+ "image_id": 7278,
1340
+ "path": "000000007278.jpg",
1341
+ "sha256": "d94baa129ccde212e3407b4032ad3bb2ceacd8191b4cabb9bd2683aa00a63284",
1342
+ "categories": [
1343
+ 0,
1344
+ 37
1345
+ ],
1346
+ "size": [
1347
+ 640,
1348
+ 482
1349
+ ]
1350
+ },
1351
+ {
1352
+ "index": 72,
1353
+ "image_id": 7281,
1354
+ "path": "000000007281.jpg",
1355
+ "sha256": "046b762eb4ffa68dc0f1c15853a71753061f622a93fcc668cc2c060a2fb26e24",
1356
+ "categories": [
1357
+ 17,
1358
+ 17,
1359
+ 0,
1360
+ 0,
1361
+ 0,
1362
+ 0,
1363
+ 0,
1364
+ 0,
1365
+ 0,
1366
+ 0,
1367
+ 0,
1368
+ 0
1369
+ ],
1370
+ "size": [
1371
+ 640,
1372
+ 361
1373
+ ]
1374
+ },
1375
+ {
1376
+ "index": 73,
1377
+ "image_id": 7386,
1378
+ "path": "000000007386.jpg",
1379
+ "sha256": "ca6798acb1c4f8604b2ff883f6d4efa9014688e1188f427da8a5c0a175e23285",
1380
+ "categories": [
1381
+ 16,
1382
+ 3,
1383
+ 1,
1384
+ 1,
1385
+ 7
1386
+ ],
1387
+ "size": [
1388
+ 600,
1389
+ 400
1390
+ ]
1391
+ },
1392
+ {
1393
+ "index": 74,
1394
+ "image_id": 7511,
1395
+ "path": "000000007511.jpg",
1396
+ "sha256": "d2f81589f5cafa0810ca36315c0a84da6286ba7412a3bb4c779309a56069b4a3",
1397
+ "categories": [
1398
+ 0,
1399
+ 0,
1400
+ 0,
1401
+ 0,
1402
+ 0,
1403
+ 33,
1404
+ 33,
1405
+ 24,
1406
+ 0,
1407
+ 0,
1408
+ 0,
1409
+ 0,
1410
+ 0,
1411
+ 0,
1412
+ 0,
1413
+ 0,
1414
+ 0
1415
+ ],
1416
+ "size": [
1417
+ 640,
1418
+ 480
1419
+ ]
1420
+ },
1421
+ {
1422
+ "index": 75,
1423
+ "image_id": 7574,
1424
+ "path": "000000007574.jpg",
1425
+ "sha256": "bfd81421e4e2e069390236672de90e9624ba503b222e54bdfea37d138de274dc",
1426
+ "categories": [
1427
+ 39,
1428
+ 39,
1429
+ 72,
1430
+ 40,
1431
+ 40,
1432
+ 69,
1433
+ 71,
1434
+ 45,
1435
+ 45,
1436
+ 68,
1437
+ 75,
1438
+ 75
1439
+ ],
1440
+ "size": [
1441
+ 640,
1442
+ 480
1443
+ ]
1444
+ },
1445
+ {
1446
+ "index": 76,
1447
+ "image_id": 7784,
1448
+ "path": "000000007784.jpg",
1449
+ "sha256": "9b3c8a668758363c40e0ce9589a8393b90c36ec1604ebc340acd50cabc5f5ee6",
1450
+ "categories": [
1451
+ 33
1452
+ ],
1453
+ "size": [
1454
+ 500,
1455
+ 375
1456
+ ]
1457
+ },
1458
+ {
1459
+ "index": 77,
1460
+ "image_id": 7795,
1461
+ "path": "000000007795.jpg",
1462
+ "sha256": "b388fc56bff44861f3b8be932ab99b532aaac1096bbe2abab24c6c02821e8775",
1463
+ "categories": [
1464
+ 59,
1465
+ 59,
1466
+ 74,
1467
+ 65
1468
+ ],
1469
+ "size": [
1470
+ 640,
1471
+ 427
1472
+ ]
1473
+ },
1474
+ {
1475
+ "index": 78,
1476
+ "image_id": 7816,
1477
+ "path": "000000007816.jpg",
1478
+ "sha256": "24f8fce4cd728856b53cfa5200e366daa529c8e378c73e1d6c7032afdfcd3dd0",
1479
+ "categories": [
1480
+ 3,
1481
+ 0,
1482
+ 0,
1483
+ 0,
1484
+ 0,
1485
+ 0,
1486
+ 0,
1487
+ 0,
1488
+ 0,
1489
+ 0,
1490
+ 0
1491
+ ],
1492
+ "size": [
1493
+ 640,
1494
+ 427
1495
+ ]
1496
+ },
1497
+ {
1498
+ "index": 79,
1499
+ "image_id": 7818,
1500
+ "path": "000000007818.jpg",
1501
+ "sha256": "f8363f92bb25e62ad8b4689c686e82cc9748c0a1d46c0e9fc1cfa5c307edce51",
1502
+ "categories": [
1503
+ 60,
1504
+ 40,
1505
+ 40,
1506
+ 40,
1507
+ 43,
1508
+ 75,
1509
+ 42,
1510
+ 56,
1511
+ 43,
1512
+ 43,
1513
+ 56
1514
+ ],
1515
+ "size": [
1516
+ 640,
1517
+ 427
1518
+ ]
1519
+ },
1520
+ {
1521
+ "index": 80,
1522
+ "image_id": 7888,
1523
+ "path": "000000007888.jpg",
1524
+ "sha256": "17b4c0e6d3f46869b6d8ffb42579c835ed299144fcc816e58f089ed6129f766d",
1525
+ "categories": [
1526
+ 74,
1527
+ 74
1528
+ ],
1529
+ "size": [
1530
+ 638,
1531
+ 640
1532
+ ]
1533
+ },
1534
+ {
1535
+ "index": 81,
1536
+ "image_id": 7977,
1537
+ "path": "000000007977.jpg",
1538
+ "sha256": "e2f6f1a54125dce42b998f65d0a0c5b5be4982cbc214a74323761c19ab703e8d",
1539
+ "categories": [
1540
+ 0,
1541
+ 36,
1542
+ 0,
1543
+ 0,
1544
+ 0,
1545
+ 0
1546
+ ],
1547
+ "size": [
1548
+ 429,
1549
+ 640
1550
+ ]
1551
+ },
1552
+ {
1553
+ "index": 82,
1554
+ "image_id": 7991,
1555
+ "path": "000000007991.jpg",
1556
+ "sha256": "fe8a9046a68e3bb62f47d2ea9c16a2bdf8c93db8f2f9216e36b9123dc435a406",
1557
+ "categories": [
1558
+ 43,
1559
+ 51,
1560
+ 43
1561
+ ],
1562
+ "size": [
1563
+ 640,
1564
+ 359
1565
+ ]
1566
+ },
1567
+ {
1568
+ "index": 83,
1569
+ "image_id": 8021,
1570
+ "path": "000000008021.jpg",
1571
+ "sha256": "9553987e9ca695ea9d68192364552fdd6cc1610925c29f405bc973727f2af0e9",
1572
+ "categories": [
1573
+ 27,
1574
+ 0,
1575
+ 0,
1576
+ 0,
1577
+ 39,
1578
+ 39,
1579
+ 39
1580
+ ],
1581
+ "size": [
1582
+ 640,
1583
+ 480
1584
+ ]
1585
+ },
1586
+ {
1587
+ "index": 84,
1588
+ "image_id": 8211,
1589
+ "path": "000000008211.jpg",
1590
+ "sha256": "102e11f581ed8b072bb87e08c0c5ef71aecf584626807808190ff911a74e2527",
1591
+ "categories": [
1592
+ 3,
1593
+ 56,
1594
+ 56,
1595
+ 0,
1596
+ 0,
1597
+ 1
1598
+ ],
1599
+ "size": [
1600
+ 640,
1601
+ 459
1602
+ ]
1603
+ },
1604
+ {
1605
+ "index": 85,
1606
+ "image_id": 8277,
1607
+ "path": "000000008277.jpg",
1608
+ "sha256": "f50009897370f60e3b661cb1d3845d87f6eecf4d80b7db913459abac86cde883",
1609
+ "categories": [
1610
+ 42,
1611
+ 45,
1612
+ 45,
1613
+ 50,
1614
+ 50,
1615
+ 50,
1616
+ 50,
1617
+ 50,
1618
+ 50,
1619
+ 50,
1620
+ 50,
1621
+ 50,
1622
+ 50
1623
+ ],
1624
+ "size": [
1625
+ 612,
1626
+ 612
1627
+ ]
1628
+ },
1629
+ {
1630
+ "index": 86,
1631
+ "image_id": 8532,
1632
+ "path": "000000008532.jpg",
1633
+ "sha256": "3e0a6d522287e931f499193bbfba59060120479a6d10f83661f474b2e8179493",
1634
+ "categories": [
1635
+ 27,
1636
+ 0
1637
+ ],
1638
+ "size": [
1639
+ 640,
1640
+ 426
1641
+ ]
1642
+ },
1643
+ {
1644
+ "index": 87,
1645
+ "image_id": 8629,
1646
+ "path": "000000008629.jpg",
1647
+ "sha256": "b85ce1980d8558d95d8685f347a7db85530c5105882eba97bbfac85afc86d268",
1648
+ "categories": [
1649
+ 42,
1650
+ 53,
1651
+ 53,
1652
+ 53,
1653
+ 53,
1654
+ 53,
1655
+ 53
1656
+ ],
1657
+ "size": [
1658
+ 640,
1659
+ 640
1660
+ ]
1661
+ },
1662
+ {
1663
+ "index": 88,
1664
+ "image_id": 8690,
1665
+ "path": "000000008690.jpg",
1666
+ "sha256": "49fd9f0199cbb9be76706dfa6338e8dfecf2a181f7304b79cba39d36b863509a",
1667
+ "categories": [
1668
+ 0,
1669
+ 0,
1670
+ 0,
1671
+ 0,
1672
+ 18
1673
+ ],
1674
+ "size": [
1675
+ 640,
1676
+ 480
1677
+ ]
1678
+ },
1679
+ {
1680
+ "index": 89,
1681
+ "image_id": 8762,
1682
+ "path": "000000008762.jpg",
1683
+ "sha256": "a8decf5853cac7e0c29e9502f5d59bf00580199e5c66b7bb08bca55c373d0cb3",
1684
+ "categories": [
1685
+ 9,
1686
+ 9,
1687
+ 9,
1688
+ 9,
1689
+ 2,
1690
+ 2,
1691
+ 9
1692
+ ],
1693
+ "size": [
1694
+ 640,
1695
+ 427
1696
+ ]
1697
+ },
1698
+ {
1699
+ "index": 90,
1700
+ "image_id": 8844,
1701
+ "path": "000000008844.jpg",
1702
+ "sha256": "837749c3cd062be1f4881d6a696af767b5b20cdabb2a989c9e79cd03d4368515",
1703
+ "categories": [
1704
+ 0,
1705
+ 46,
1706
+ 46,
1707
+ 46,
1708
+ 46,
1709
+ 0,
1710
+ 0
1711
+ ],
1712
+ "size": [
1713
+ 640,
1714
+ 426
1715
+ ]
1716
+ },
1717
+ {
1718
+ "index": 91,
1719
+ "image_id": 8899,
1720
+ "path": "000000008899.jpg",
1721
+ "sha256": "01f50fb2d84379f3458ae1891cf161be6f4853f5d654ddc999b576c8be533501",
1722
+ "categories": [
1723
+ 10,
1724
+ 1
1725
+ ],
1726
+ "size": [
1727
+ 640,
1728
+ 539
1729
+ ]
1730
+ },
1731
+ {
1732
+ "index": 92,
1733
+ "image_id": 9378,
1734
+ "path": "000000009378.jpg",
1735
+ "sha256": "788f6b877c0493f1f665eabb77be1c2b3e620b2ba8258b1b78d4239f0d224a12",
1736
+ "categories": [
1737
+ 0,
1738
+ 0,
1739
+ 0,
1740
+ 0,
1741
+ 29,
1742
+ 0,
1743
+ 0,
1744
+ 0,
1745
+ 0,
1746
+ 0
1747
+ ],
1748
+ "size": [
1749
+ 600,
1750
+ 400
1751
+ ]
1752
+ },
1753
+ {
1754
+ "index": 93,
1755
+ "image_id": 9400,
1756
+ "path": "000000009400.jpg",
1757
+ "sha256": "187c9f45b30ba9cf88e006868c8d34aee42a49faffeb5339f400243a6eafda64",
1758
+ "categories": [
1759
+ 0,
1760
+ 0,
1761
+ 0,
1762
+ 63,
1763
+ 63,
1764
+ 63,
1765
+ 63,
1766
+ 64,
1767
+ 66,
1768
+ 0,
1769
+ 0,
1770
+ 0,
1771
+ 0,
1772
+ 0,
1773
+ 41,
1774
+ 63,
1775
+ 66,
1776
+ 66,
1777
+ 66,
1778
+ 41,
1779
+ 41,
1780
+ 41,
1781
+ 66,
1782
+ 0
1783
+ ],
1784
+ "size": [
1785
+ 640,
1786
+ 480
1787
+ ]
1788
+ },
1789
+ {
1790
+ "index": 94,
1791
+ "image_id": 9448,
1792
+ "path": "000000009448.jpg",
1793
+ "sha256": "e09ae0fbc067abe41e59e2e2bc68e0fbca0f8e296919c8e5721b0f7c5046d624",
1794
+ "categories": [
1795
+ 25,
1796
+ 0
1797
+ ],
1798
+ "size": [
1799
+ 551,
1800
+ 640
1801
+ ]
1802
+ },
1803
+ {
1804
+ "index": 95,
1805
+ "image_id": 9483,
1806
+ "path": "000000009483.jpg",
1807
+ "sha256": "713f1ab765fa57b825473055407f259d75c2a0dd379dee78b330b7e54e653304",
1808
+ "categories": [
1809
+ 62,
1810
+ 0,
1811
+ 64,
1812
+ 66,
1813
+ 62,
1814
+ 62
1815
+ ],
1816
+ "size": [
1817
+ 640,
1818
+ 480
1819
+ ]
1820
+ },
1821
+ {
1822
+ "index": 96,
1823
+ "image_id": 9590,
1824
+ "path": "000000009590.jpg",
1825
+ "sha256": "58ca4867198c5c4745e748720e2abd67a51d0fef89ae01da6f6f584c50e9297d",
1826
+ "categories": [
1827
+ 39,
1828
+ 74,
1829
+ 60,
1830
+ 0,
1831
+ 0,
1832
+ 0,
1833
+ 41,
1834
+ 41,
1835
+ 44,
1836
+ 44,
1837
+ 45,
1838
+ 45,
1839
+ 45,
1840
+ 45,
1841
+ 45,
1842
+ 0,
1843
+ 41,
1844
+ 41,
1845
+ 41,
1846
+ 45,
1847
+ 56,
1848
+ 0,
1849
+ 41,
1850
+ 41,
1851
+ 41,
1852
+ 41,
1853
+ 44,
1854
+ 0,
1855
+ 41
1856
+ ],
1857
+ "size": [
1858
+ 640,
1859
+ 427
1860
+ ]
1861
+ },
1862
+ {
1863
+ "index": 97,
1864
+ "image_id": 9769,
1865
+ "path": "000000009769.jpg",
1866
+ "sha256": "87757b4774353e2adcf999e4446904a5ddd97bca193595266e98d46e721c575d",
1867
+ "categories": [
1868
+ 0,
1869
+ 10,
1870
+ 0,
1871
+ 0,
1872
+ 7
1873
+ ],
1874
+ "size": [
1875
+ 640,
1876
+ 480
1877
+ ]
1878
+ },
1879
+ {
1880
+ "index": 98,
1881
+ "image_id": 9772,
1882
+ "path": "000000009772.jpg",
1883
+ "sha256": "399bfa699b513e0c5071cdf1dcacbe599fb0c1fdfd15c67cb22051d1860b0575",
1884
+ "categories": [
1885
+ 62,
1886
+ 0,
1887
+ 71,
1888
+ 71
1889
+ ],
1890
+ "size": [
1891
+ 550,
1892
+ 640
1893
+ ]
1894
+ },
1895
+ {
1896
+ "index": 99,
1897
+ "image_id": 9891,
1898
+ "path": "000000009891.jpg",
1899
+ "sha256": "9e1b89352fe976f8fb586594a526c8dab8a4d98aa5bdf5d5813bca80d1cd5b14",
1900
+ "categories": [
1901
+ 27,
1902
+ 2,
1903
+ 0,
1904
+ 0,
1905
+ 0,
1906
+ 0,
1907
+ 28,
1908
+ 2,
1909
+ 2,
1910
+ 24,
1911
+ 28,
1912
+ 24,
1913
+ 2
1914
+ ],
1915
+ "size": [
1916
+ 640,
1917
+ 480
1918
+ ]
1919
+ }
1920
+ ]
1921
+ }
reports/round2/coco_revision.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"repo": "detection-datasets/coco", "parquet_revision": "26ddc382fe75dfc2a0655b5977e296ea10efebce"}
reports/round2/coco_validation200.json ADDED
The diff for this file is too large to render. See raw diff
 
reports/round2/environment-lock.txt ADDED
@@ -0,0 +1,60 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ aiohappyeyeballs==2.7.1
2
+ aiohttp==3.14.3
3
+ aiosignal==1.4.0
4
+ anyio==4.15.1
5
+ attrs==26.1.0
6
+ certifi==2026.7.22
7
+ charset-normalizer==3.5.1
8
+ click==8.5.0
9
+ datasets==5.0.1
10
+ dill==0.4.1
11
+ filelock==3.32.3
12
+ flatbuffers==25.12.19
13
+ frozenlist==1.8.0
14
+ fsspec==2026.6.0
15
+ h11==0.16.0
16
+ hf-xet==1.6.0
17
+ httpcore==1.0.9
18
+ httpx==0.28.1
19
+ huggingface-hub==1.32.0
20
+ idna==3.20
21
+ imageio==2.37.4
22
+ iniconfig==2.3.0
23
+ jinja2==3.1.6
24
+ lazy-loader==0.5
25
+ markupsafe==3.0.3
26
+ ml-dtypes==0.6.0
27
+ mpmath==1.3.0
28
+ multidict==6.9.0
29
+ multiprocess==0.70.19
30
+ networkx==3.6.1
31
+ numpy==2.5.2
32
+ onnx==1.23.0
33
+ onnxruntime==1.30.0
34
+ packaging==26.3
35
+ pandas==3.0.6
36
+ pillow==12.3.0
37
+ pluggy==1.6.0
38
+ propcache==0.5.4
39
+ protobuf==7.36.2
40
+ pyarrow==25.0.1
41
+ pygments==2.21.0
42
+ pytest==9.1.1
43
+ python-dateutil==2.9.0.post0
44
+ pyyaml==6.0.3
45
+ requests==2.34.2
46
+ safetensors==0.8.0
47
+ scikit-image==0.26.0
48
+ scipy==1.18.1
49
+ setuptools==78.1.0
50
+ six==1.17.0
51
+ socksio==1.0.0
52
+ sympy==1.14.0
53
+ tifffile==2026.9.15
54
+ torch==2.14.0+cpu
55
+ torchvision==0.29.0+cpu
56
+ tqdm==4.70.1
57
+ typing-extensions==4.16.0
58
+ urllib3==2.8.0
59
+ xxhash==4.0.1
60
+ yarl==1.25.1
reports/round2/final_coco100.json ADDED
The diff for this file is too large to render. See raw diff
 
reports/round2/final_coco100.png ADDED

Git LFS Details

  • SHA256: 2ba6c18c60e7858574f45f1fe6b1d25ef14031d8b92a3986ec057f5c945c337e
  • Pointer size: 132 Bytes
  • Size of remote file: 1.3 MB
reports/round2/final_integrity.json ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "image_wrapper_checks": [
3
+ {
4
+ "name": "astronaut",
5
+ "shape": [
6
+ 512,
7
+ 512,
8
+ 3
9
+ ],
10
+ "max_rgb8_difference": 1.0,
11
+ "mean_rgb8_difference": 2.7974446614583332e-05
12
+ },
13
+ {
14
+ "name": "coffee",
15
+ "shape": [
16
+ 400,
17
+ 600,
18
+ 3
19
+ ],
20
+ "max_rgb8_difference": 1.0,
21
+ "mean_rgb8_difference": 0.00011805555555555556
22
+ },
23
+ {
24
+ "name": "chelsea",
25
+ "shape": [
26
+ 300,
27
+ 451,
28
+ 3
29
+ ],
30
+ "max_rgb8_difference": 1.0,
31
+ "mean_rgb8_difference": 1.7245627001724563e-05
32
+ }
33
+ ],
34
+ "parameters": 3968892,
35
+ "changed_learned_tensors": 65,
36
+ "bin_centers_identical_to_audited_baseline": true,
37
+ "model_sha256": "0e4c417375684a044860f8af3ac3a2fb44e1a5729254ca33abda758ea71aea6e",
38
+ "onnx_sha256": "0ef86749901e66ad53b1e8e1d940330572e2f4c9347ac01e7bc02c7683f8c79a"
39
+ }
reports/round2/final_results.json ADDED
@@ -0,0 +1,207 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "imagenette200": {
3
+ "comparisons": {
4
+ "previous_raw_to_new_raw": {
5
+ "n": 200,
6
+ "ab_error_before": 13.317798409461975,
7
+ "ab_error_after": 13.055905781388283,
8
+ "relative_error_reduction_percent": 1.966485901210402,
9
+ "mean_paired_improvement": 0.26189262807369235,
10
+ "bootstrap_95ci": [
11
+ 0.09549147181212903,
12
+ 0.4279274906665086
13
+ ],
14
+ "images_improved": 127,
15
+ "fine_excess_edge_reduction_percent": 3.833467907331034,
16
+ "coarse_excess_edge_reduction_percent": 9.867497214465548
17
+ },
18
+ "previous_raw_to_new_guided8": {
19
+ "n": 200,
20
+ "ab_error_before": 13.317798409461975,
21
+ "ab_error_after": 12.846773383617402,
22
+ "relative_error_reduction_percent": 3.536808497641186,
23
+ "mean_paired_improvement": 0.47102502584457395,
24
+ "bootstrap_95ci": [
25
+ 0.3075495718717575,
26
+ 0.63933595174551
27
+ ],
28
+ "images_improved": 142,
29
+ "fine_excess_edge_reduction_percent": 85.82711416487938,
30
+ "coarse_excess_edge_reduction_percent": 55.24418461306793
31
+ },
32
+ "previous_guided8_to_new_guided8": {
33
+ "n": 200,
34
+ "ab_error_before": 13.076925607919692,
35
+ "ab_error_after": 12.846773383617402,
36
+ "relative_error_reduction_percent": 1.759987256965856,
37
+ "mean_paired_improvement": 0.23015222430229187,
38
+ "bootstrap_95ci": [
39
+ 0.060473272860050206,
40
+ 0.3983106223940849
41
+ ],
42
+ "images_improved": 123,
43
+ "fine_excess_edge_reduction_percent": 7.75156130633905,
44
+ "coarse_excess_edge_reduction_percent": 9.983848905802084
45
+ }
46
+ },
47
+ "summary": {
48
+ "previous_raw": {
49
+ "ce": 2.9057628363370895,
50
+ "ab_error": 13.317798409461975,
51
+ "chroma": 15.02951672077179,
52
+ "target_chroma": 15.572654025255469,
53
+ "excess_chroma_edge": 0.5394541963189841,
54
+ "n": 200,
55
+ "chroma_ratio": 0.9651223674780915,
56
+ "coarse_excess_edge": 1.1962259417027234,
57
+ "false_color_jump_fraction": 0.0018527497181594298,
58
+ "low_chroma_sources": 12,
59
+ "color_source_ab_error": 13.674592086609374
60
+ },
61
+ "previous_guided8": {
62
+ "ce": 2.9057628363370895,
63
+ "ab_error": 13.076925607919692,
64
+ "chroma": 14.679913553595544,
65
+ "target_chroma": 15.572654025255469,
66
+ "excess_chroma_edge": 0.08288078200537712,
67
+ "n": 200,
68
+ "chroma_ratio": 0.9426725547095509,
69
+ "coarse_excess_edge": 0.594760681912303,
70
+ "false_color_jump_fraction": 0.0,
71
+ "low_chroma_sources": 12,
72
+ "color_source_ab_error": 13.434583253048835
73
+ },
74
+ "new_raw": {
75
+ "ce": 2.900379472374916,
76
+ "ab_error": 13.055905781388283,
77
+ "chroma": 14.313646669387817,
78
+ "target_chroma": 15.572654025255469,
79
+ "excess_chroma_edge": 0.5187743928283453,
80
+ "n": 200,
81
+ "chroma_ratio": 0.9191526791884148,
82
+ "coarse_excess_edge": 1.0781883802264929,
83
+ "false_color_jump_fraction": 0.001551891862541197,
84
+ "low_chroma_sources": 12,
85
+ "color_source_ab_error": 13.417864246571318,
86
+ "low_chroma_threshold": 3
87
+ },
88
+ "new_guided8": {
89
+ "ce": 2.900379472374916,
90
+ "ab_error": 12.846773383617402,
91
+ "chroma": 14.004629806876183,
92
+ "target_chroma": 15.572654025255469,
93
+ "excess_chroma_edge": 0.07645622737705708,
94
+ "n": 200,
95
+ "chroma_ratio": 0.8993091212431554,
96
+ "coarse_excess_edge": 0.5353806740790605,
97
+ "false_color_jump_fraction": 0.0,
98
+ "low_chroma_sources": 12,
99
+ "color_source_ab_error": 13.210732571622158,
100
+ "low_chroma_threshold": 3
101
+ }
102
+ }
103
+ },
104
+ "coco100": {
105
+ "comparisons": {
106
+ "previous_raw_to_new_raw": {
107
+ "n": 100,
108
+ "ab_error_before": 15.266220955848693,
109
+ "ab_error_after": 14.770637459754944,
110
+ "relative_error_reduction_percent": 3.246274880515765,
111
+ "mean_paired_improvement": 0.49558349609375,
112
+ "bootstrap_95ci": [
113
+ 0.12661694943904878,
114
+ 0.8486135826110839
115
+ ],
116
+ "images_improved": 66,
117
+ "fine_excess_edge_reduction_percent": 1.40337050048579,
118
+ "coarse_excess_edge_reduction_percent": 9.182688028442687
119
+ },
120
+ "previous_raw_to_new_guided8": {
121
+ "n": 100,
122
+ "ab_error_before": 15.266220955848693,
123
+ "ab_error_after": 14.531693441867828,
124
+ "relative_error_reduction_percent": 4.811456064373665,
125
+ "mean_paired_improvement": 0.7345275139808655,
126
+ "bootstrap_95ci": [
127
+ 0.3619160271286964,
128
+ 1.0888569805026054
129
+ ],
130
+ "images_improved": 76,
131
+ "fine_excess_edge_reduction_percent": 84.3489471902113,
132
+ "coarse_excess_edge_reduction_percent": 53.952045348000254
133
+ },
134
+ "previous_guided8_to_new_guided8": {
135
+ "n": 100,
136
+ "ab_error_before": 14.992383620738984,
137
+ "ab_error_after": 14.531693441867828,
138
+ "relative_error_reduction_percent": 3.0728281140957603,
139
+ "mean_paired_improvement": 0.46069017887115477,
140
+ "bootstrap_95ci": [
141
+ 0.09070354998111727,
142
+ 0.8119295918345449
143
+ ],
144
+ "images_improved": 65,
145
+ "fine_excess_edge_reduction_percent": 6.799819678037888,
146
+ "coarse_excess_edge_reduction_percent": 10.05219053918135
147
+ }
148
+ },
149
+ "summary": {
150
+ "previous_raw": {
151
+ "ce": 3.095473835468292,
152
+ "ab_error": 15.266220955848693,
153
+ "chroma": 14.049296131134033,
154
+ "target_chroma": 14.326628150648903,
155
+ "excess_chroma_edge": 0.5895961074531079,
156
+ "n": 100,
157
+ "chroma_ratio": 0.9806421988064018,
158
+ "coarse_excess_edge": 1.4608290269970894,
159
+ "false_color_jump_fraction": 0.001828507996387998,
160
+ "low_chroma_sources": 12,
161
+ "color_source_ab_error": 15.702739195390182
162
+ },
163
+ "previous_guided8": {
164
+ "ce": 3.095473835468292,
165
+ "ab_error": 14.992383620738984,
166
+ "chroma": 13.739825406074523,
167
+ "target_chroma": 14.326628150648903,
168
+ "excess_chroma_edge": 0.09901053605601191,
169
+ "n": 100,
170
+ "chroma_ratio": 0.9590411129259468,
171
+ "coarse_excess_edge": 0.7478579989075661,
172
+ "false_color_jump_fraction": 0.0,
173
+ "low_chroma_sources": 12,
174
+ "color_source_ab_error": 15.42814136093313
175
+ },
176
+ "new_raw": {
177
+ "ce": 3.0355803537368775,
178
+ "ab_error": 14.770637459754944,
179
+ "chroma": 14.113198335170745,
180
+ "target_chroma": 14.326628150648903,
181
+ "excess_chroma_edge": 0.5813218896090985,
182
+ "n": 100,
183
+ "chroma_ratio": 0.9851025786923568,
184
+ "coarse_excess_edge": 1.3266856548190118,
185
+ "false_color_jump_fraction": 0.0016237745340367837,
186
+ "low_chroma_sources": 12,
187
+ "color_source_ab_error": 15.167718676003544,
188
+ "low_chroma_threshold": 3
189
+ },
190
+ "new_guided8": {
191
+ "ce": 3.0355803537368775,
192
+ "ab_error": 14.531693441867828,
193
+ "chroma": 13.846259930133819,
194
+ "target_chroma": 14.326628150648903,
195
+ "excess_chroma_edge": 0.09227799814194441,
196
+ "n": 100,
197
+ "chroma_ratio": 0.9664702527723995,
198
+ "coarse_excess_edge": 0.6726818878948688,
199
+ "false_color_jump_fraction": 0.0,
200
+ "low_chroma_sources": 12,
201
+ "color_source_ab_error": 14.927411946383389,
202
+ "low_chroma_threshold": 3
203
+ }
204
+ }
205
+ },
206
+ "caveat": "Conditional image-bootstrap intervals; convenience COCO sample; artifact proxies are not blotch counts. No test-based retuning."
207
+ }
reports/round2/final_test200.json ADDED
The diff for this file is too large to render. See raw diff
 
reports/round2/final_test_comparison.png ADDED

Git LFS Details

  • SHA256: 9681e45700aa5def0f69f67a49a45cadfddb06c1dc4b321995e09641eebf7f73
  • Pointer size: 132 Bytes
  • Size of remote file: 1.73 MB
reports/round2/final_visual.json ADDED
@@ -0,0 +1,73 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "release-v2",
3
+ "results": {
4
+ "astronaut": {
5
+ "Grayscale": {
6
+ "ab_error": 20.617759704589844
7
+ },
8
+ "Previous repair": {
9
+ "ab_error": 14.877237319946289
10
+ },
11
+ "New weights, raw": {
12
+ "ab_error": 14.900100708007812
13
+ },
14
+ "New weights + smoothing": {
15
+ "ab_error": 14.482976913452148
16
+ },
17
+ "Reference": {
18
+ "ab_error": 0.0
19
+ }
20
+ },
21
+ "coffee": {
22
+ "Grayscale": {
23
+ "ab_error": 43.01523208618164
24
+ },
25
+ "Previous repair": {
26
+ "ab_error": 26.923748016357422
27
+ },
28
+ "New weights, raw": {
29
+ "ab_error": 21.989187240600586
30
+ },
31
+ "New weights + smoothing": {
32
+ "ab_error": 20.511058807373047
33
+ },
34
+ "Reference": {
35
+ "ab_error": 0.0
36
+ }
37
+ },
38
+ "chelsea": {
39
+ "Grayscale": {
40
+ "ab_error": 22.89802360534668
41
+ },
42
+ "Previous repair": {
43
+ "ab_error": 22.859899520874023
44
+ },
45
+ "New weights, raw": {
46
+ "ab_error": 18.394851684570312
47
+ },
48
+ "New weights + smoothing": {
49
+ "ab_error": 18.324064254760742
50
+ },
51
+ "Reference": {
52
+ "ab_error": 0.0
53
+ }
54
+ },
55
+ "rocket": {
56
+ "Grayscale": {
57
+ "ab_error": 18.316810607910156
58
+ },
59
+ "Previous repair": {
60
+ "ab_error": 21.93663215637207
61
+ },
62
+ "New weights, raw": {
63
+ "ab_error": 23.127243041992188
64
+ },
65
+ "New weights + smoothing": {
66
+ "ab_error": 22.861391067504883
67
+ },
68
+ "Reference": {
69
+ "ab_error": 0.0
70
+ }
71
+ }
72
+ }
73
+ }
reports/round2/final_visual.png ADDED

Git LFS Details

  • SHA256: df2b3c271150059ab2f0d0a10f7fe889f1d3878c6096e14e0fd5df13edd7022a
  • Pointer size: 132 Bytes
  • Size of remote file: 1.23 MB
reports/round2/hub_refs.json ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ {
2
+ "main": "6c47ea40724d8fcd67d4f36ce837dc1cb5b1b2a8",
3
+ "stable": "b875f0fd5ff7f2e39f2c6068a60b268fb35586cc"
4
+ }
reports/round2/jobs.json ADDED
@@ -0,0 +1,46 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "new_custom_plugin_tools_available": false,
3
+ "connected_plugin": "Hugging Face",
4
+ "connected_oauth_scopes": [
5
+ "jobs",
6
+ "openid",
7
+ "profile",
8
+ "read-mcp",
9
+ "read-repos"
10
+ ],
11
+ "artifact_transport_preflight": {
12
+ "job_id": "6aaf8a1851992417dfccc4a3",
13
+ "bytes": 17825792,
14
+ "sha256": "cbff78f490321c640b69d664174fa5d9880b230aa0dff9061965c214cfdf3c57",
15
+ "verified_local_roundtrip": true
16
+ },
17
+ "fp16_startup_failures": [
18
+ "6aaf8afc52d0dbd7f1d72ffd",
19
+ "6aaf8b0851992417dfccc4db"
20
+ ],
21
+ "bf16_jobs": {
22
+ "classification": "6aaf8ceb52d0dbd7f1d73094",
23
+ "multiscale": "6aaf8cf752d0dbd7f1d7309a",
24
+ "mixed_data": "6aaf8dad51992417dfccc557"
25
+ },
26
+ "hardware": "l4x1",
27
+ "hourly_price_usd_at_submission": 0.8,
28
+ "timeout_per_job_minutes": 60,
29
+ "pricing_source": "https://huggingface.co/docs/hub/jobs-pricing",
30
+ "recovery_jobs": {
31
+ "classification": "6aaf900452d0dbd7f1d7313c",
32
+ "mixed_last": "6aaf912f52d0dbd7f1d7319e",
33
+ "mixed_best": "6aaf91b151992417dfccc5f1",
34
+ "multiscale": "6aaf913a51992417dfccc5de",
35
+ "validation_images": "6aaf925252d0dbd7f1d731d5"
36
+ },
37
+ "failed_recovery_jobs_default_log_truncation": [
38
+ "6aaf90be51992417dfccc5cc",
39
+ "6aaf90fa52d0dbd7f1d7318a"
40
+ ],
41
+ "final_gpu_statuses": {
42
+ "classification": "COMPLETED",
43
+ "multiscale": "COMPLETED",
44
+ "mixed_data": "COMPLETED"
45
+ }
46
+ }
reports/round2/manifest.json ADDED
The diff for this file is too large to render. See raw diff
 
reports/round2/onnx_benchmark.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "runtime": "1.30.0",
3
+ "provider": "CPUExecutionProvider",
4
+ "threads": 2,
5
+ "shape": [
6
+ 1,
7
+ 1,
8
+ 256,
9
+ 256
10
+ ],
11
+ "warmup": 5,
12
+ "iterations": 20,
13
+ "median_ms": 908.7965385001553,
14
+ "p90_ms": 1238.8696021996116,
15
+ "peak_process_rss_mib": 507.921875,
16
+ "platform": "Linux-6.18.44-x86_64-with-glibc2.39",
17
+ "scope": "Network,decode,guided8 only; excludes file loading and full-resolution color conversions. Shared host measurement, not deployment SLA."
18
+ }
reports/round2/onnx_export.json ADDED
@@ -0,0 +1,54 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "input": "N,1,H,W Lab L*/50-1; H,W>=8",
3
+ "output": "N,2,H,W Lab chroma a,b",
4
+ "guided_radius": 8,
5
+ "guided_epsilon": 0.001,
6
+ "temperature": 0.38,
7
+ "opset": 17,
8
+ "checks": [
9
+ {
10
+ "shape": [
11
+ 1,
12
+ 1,
13
+ 256,
14
+ 256
15
+ ],
16
+ "max_ab_difference": 8.58306884765625e-06,
17
+ "mean_ab_difference": 1.0279118214384653e-06
18
+ },
19
+ {
20
+ "shape": [
21
+ 1,
22
+ 1,
23
+ 173,
24
+ 241
25
+ ],
26
+ "max_ab_difference": 7.3909759521484375e-06,
27
+ "mean_ab_difference": 1.0783561492644367e-06
28
+ },
29
+ {
30
+ "shape": [
31
+ 2,
32
+ 1,
33
+ 64,
34
+ 80
35
+ ],
36
+ "max_ab_difference": 7.152557373046875e-06,
37
+ "mean_ab_difference": 9.707116532808868e-07
38
+ },
39
+ {
40
+ "shape": [
41
+ 1,
42
+ 1,
43
+ 8,
44
+ 9
45
+ ],
46
+ "max_ab_difference": 1.5497207641601562e-06,
47
+ "mean_ab_difference": 3.6218099808138504e-07
48
+ }
49
+ ],
50
+ "parameters": 3968892,
51
+ "weights_source": "release-v2",
52
+ "runtime": "1.30.0",
53
+ "bytes": 15885757
54
+ }
reports/round2/selection_protocol.md ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Round 2 frozen evaluation protocol
2
+
3
+ The 13-way spatial ablation used only the existing 50-image validation set.
4
+ Before scoring the new 200-image Imagenette subset or the first100 COCO-val
5
+ images, nominate guided8 for a single-pass balanced option; flip_guided4 as
6
+ a two-pass fidelity option; guided16 as a stronger smoothing comparison.
7
+ The latter visibly removes some small legitimate colour details, so is not
8
+ automatically selected by its lower artifact proxy. Baseline remains in all
9
+ comparisons. Temperature stays0.38 and saturation stays1.0. No test-based
10
+ parameter retuning is planned.
11
+
12
+ Training candidates are compared on the original validation set, with raw
13
+ CE, chroma error, mean chroma, excess colour edges and visual checks.
14
+ Do not promote a candidate solely because it is grayer or smoother. Select
15
+ weights on validation before scoring them on the new200 and COCO probe.
16
+ The first100 COCO-val rows are an external convenience sample, not a
17
+ representative population benchmark. No COCO photo is used for training.
18
+
19
+ CPU pilots:400updates,batch2,LR2e-6,seed0. GPU arms:1122updates,batch32,
20
+ three full epochs,LR1e-5,seed123. Both pairs use frozen BatchNorm, fixed
21
+ original class weights and bins,11943filtered training images. Within each
22
+ pair the sole change is multiscale chroma-gradient loss coefficient0 vs10.
23
+ CPU-to-GPU comparisons also change batch,LR,seed,precision and compute;
24
+ do not attribute their differences to one factor.
25
+
26
+ ## Broader-data addendum
27
+
28
+ The mixed-data arm was added during the exploratory round. It retains the same
29
+ 1122updates and32-image batch as the classification control, adding5382 eligible
30
+ COCO training photos from two pinned shards. This is approximately2.07epochs
31
+ over17325photos, not three epochs. Its held-out200COCO training photos are
32
+ excluded from all training and are separate from the100COCO-val external probe.
33
+
34
+ For final weight selection, compare the eligible candidates on the original50
35
+ Imagenette validation photos and those200COCO training-validation photos. Use
36
+ equal-domain mean relative ab error with guided8, retaining at least90percent
37
+ of the repaired baseline mean chroma in each validation domain and checking
38
+ that artifact proxies do not regress. This is a pragmatic development gate,
39
+ not a perceptual realism metric. The100COCO-val images are not used for this
40
+ weight ranking. Inspect visual failures before accepting the winner.
reports/round2/tests-final.txt ADDED
@@ -0,0 +1,2 @@
 
 
 
1
+ .......... [100%]
2
+ 10 passed in 3.98s
reports/round2/weight_selection.json ADDED
@@ -0,0 +1,211 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "winner": "gpu_mixed_best",
3
+ "source": "runs/round2/gpu-mixed-best/run/best",
4
+ "ranking": [
5
+ {
6
+ "name": "cpu_spatial",
7
+ "source": "runs/round2/multiscale/last",
8
+ "equal_domain_relative_error": 0.953154119425918,
9
+ "chroma_retention": [
10
+ 0.9068765792010487,
11
+ 0.869829814358865
12
+ ],
13
+ "eligible": false,
14
+ "imagenette_validation": {
15
+ "ce": 2.791827552318573,
16
+ "ab_error": 12.044532117843628,
17
+ "chroma": 13.000842504501342,
18
+ "target_chroma": 14.378198699951172,
19
+ "excess_chroma_edge": 0.06385781953111291,
20
+ "n": 50,
21
+ "chroma_ratio": 0.9042052329229178,
22
+ "coarse_excess_edge": 0.4342093861103058,
23
+ "false_color_jump_fraction": 0.0,
24
+ "low_chroma_sources": 2,
25
+ "color_source_ab_error": 12.313195079565048
26
+ },
27
+ "coco_validation": {
28
+ "ce": 3.0519875353574752,
29
+ "ab_error": 13.913508603572845,
30
+ "chroma": 10.947579425573348,
31
+ "target_chroma": 14.777132248296402,
32
+ "excess_chroma_edge": 0.07253436714410783,
33
+ "n": 200,
34
+ "chroma_ratio": 0.7408460073053384,
35
+ "coarse_excess_edge": 0.45769357804208993
36
+ }
37
+ },
38
+ {
39
+ "name": "gpu_spatial",
40
+ "source": "runs/round2/gpu-multiscale/run/last",
41
+ "equal_domain_relative_error": 0.9684214344446266,
42
+ "chroma_retention": [
43
+ 0.9275792097797178,
44
+ 0.8901436032531378
45
+ ],
46
+ "eligible": false,
47
+ "imagenette_validation": {
48
+ "ce": 2.8091157150268553,
49
+ "ab_error": 12.23028130531311,
50
+ "chroma": 13.297632217407227,
51
+ "target_chroma": 14.378198699951172,
52
+ "excess_chroma_edge": 0.06588278515264392,
53
+ "n": 50,
54
+ "chroma_ratio": 0.9248468806772288,
55
+ "coarse_excess_edge": 0.4571388277411461,
56
+ "false_color_jump_fraction": 0.0,
57
+ "low_chroma_sources": 2,
58
+ "color_source_ab_error": 12.509084890286127
59
+ },
60
+ "coco_validation": {
61
+ "ce": 3.0698557448387147,
62
+ "ab_error": 14.144731167554855,
63
+ "chroma": 11.20324646949768,
64
+ "target_chroma": 14.777132248296402,
65
+ "excess_chroma_edge": 0.07468180949334055,
66
+ "n": 200,
67
+ "chroma_ratio": 0.7581475404870427,
68
+ "coarse_excess_edge": 0.48290088372305034
69
+ }
70
+ },
71
+ {
72
+ "name": "gpu_mixed_best",
73
+ "source": "runs/round2/gpu-mixed-best/run/best",
74
+ "equal_domain_relative_error": 0.9733553425967074,
75
+ "chroma_retention": [
76
+ 0.9646619475606313,
77
+ 0.9791556597287809
78
+ ],
79
+ "eligible": true,
80
+ "imagenette_validation": {
81
+ "ce": 2.844368829727173,
82
+ "ab_error": 12.428691778182984,
83
+ "chroma": 13.829244616031646,
84
+ "target_chroma": 14.378198699951172,
85
+ "excess_chroma_edge": 0.07477423369884491,
86
+ "n": 50,
87
+ "chroma_ratio": 0.9618203854755888,
88
+ "coarse_excess_edge": 0.5252502923458815,
89
+ "false_color_jump_fraction": 0.0,
90
+ "low_chroma_sources": 2,
91
+ "color_source_ab_error": 12.718736380338669
92
+ },
93
+ "coco_validation": {
94
+ "ce": 3.0475806260108946,
95
+ "ab_error": 14.058236581087112,
96
+ "chroma": 12.323542120456695,
97
+ "target_chroma": 14.777132248296402,
98
+ "excess_chroma_edge": 0.08515535289887338,
99
+ "n": 200,
100
+ "chroma_ratio": 0.8339603323153197,
101
+ "coarse_excess_edge": 0.5545959075912833
102
+ }
103
+ },
104
+ {
105
+ "name": "gpu_mixed_last",
106
+ "source": "runs/round2/gpu-mixed/run/last",
107
+ "equal_domain_relative_error": 0.9881481747719179,
108
+ "chroma_retention": [
109
+ 1.0073343740507403,
110
+ 1.043760292635895
111
+ ],
112
+ "eligible": true,
113
+ "imagenette_validation": {
114
+ "ce": 2.857785303592682,
115
+ "ab_error": 12.591957387924195,
116
+ "chroma": 14.440989928245545,
117
+ "target_chroma": 14.378198699951172,
118
+ "excess_chroma_edge": 0.07564899820834398,
119
+ "n": 50,
120
+ "chroma_ratio": 1.004367113684038,
121
+ "coarse_excess_edge": 0.5223322550207377,
122
+ "false_color_jump_fraction": 0.0,
123
+ "low_chroma_sources": 2,
124
+ "color_source_ab_error": 12.864318639039993
125
+ },
126
+ "coco_validation": {
127
+ "ce": 3.0621814727783203,
128
+ "ab_error": 14.301741573810578,
129
+ "chroma": 13.136648705601692,
130
+ "target_chroma": 14.777132248296402,
131
+ "excess_chroma_edge": 0.08942741602659225,
132
+ "n": 200,
133
+ "chroma_ratio": 0.8889849860493848,
134
+ "coarse_excess_edge": 0.5829145759344101
135
+ }
136
+ },
137
+ {
138
+ "name": "gpu_ce",
139
+ "source": "runs/round2/gpu-ce/run/last",
140
+ "equal_domain_relative_error": 0.993647719923589,
141
+ "chroma_retention": [
142
+ 0.9771835119810507,
143
+ 0.9507257987531513
144
+ ],
145
+ "eligible": true,
146
+ "imagenette_validation": {
147
+ "ce": 2.829431531429291,
148
+ "ab_error": 12.570073976516724,
149
+ "chroma": 14.008751828670501,
150
+ "target_chroma": 14.378198699951172,
151
+ "excess_chroma_edge": 0.07572143979370594,
152
+ "n": 50,
153
+ "chroma_ratio": 0.9743050656768344,
154
+ "coarse_excess_edge": 0.5403172910213471,
155
+ "false_color_jump_fraction": 0.0,
156
+ "low_chroma_sources": 2,
157
+ "color_source_ab_error": 12.854205856720606
158
+ },
159
+ "coco_validation": {
160
+ "ce": 3.0947927451133728,
161
+ "ab_error": 14.488478088378907,
162
+ "chroma": 11.965727113485336,
163
+ "target_chroma": 14.777132248296402,
164
+ "excess_chroma_edge": 0.08773526596371084,
165
+ "n": 200,
166
+ "chroma_ratio": 0.8097462289995285,
167
+ "coarse_excess_edge": 0.5938316452503204
168
+ }
169
+ },
170
+ {
171
+ "name": "repaired",
172
+ "source": "checkpoints/main-repaired",
173
+ "equal_domain_relative_error": 1.0,
174
+ "chroma_retention": [
175
+ 1.0,
176
+ 1.0
177
+ ],
178
+ "eligible": true,
179
+ "imagenette_validation": {
180
+ "ce": 2.844365568161011,
181
+ "ab_error": 12.58308430671692,
182
+ "chroma": 14.33584547519684,
183
+ "target_chroma": 14.378198699951172,
184
+ "excess_chroma_edge": 0.0793773028627038,
185
+ "n": 50,
186
+ "chroma_ratio": 0.9970543441749433,
187
+ "coarse_excess_edge": 0.5623898504674435,
188
+ "false_color_jump_fraction": 0.0,
189
+ "low_chroma_sources": 2,
190
+ "color_source_ab_error": 12.839716345071793
191
+ },
192
+ "coco_validation": {
193
+ "ce": 3.099011971950531,
194
+ "ab_error": 14.6595640873909,
195
+ "chroma": 12.585886623859405,
196
+ "target_chroma": 14.777132248296402,
197
+ "excess_chroma_edge": 0.09364423044025898,
198
+ "n": 200,
199
+ "chroma_ratio": 0.851713743396347,
200
+ "coarse_excess_edge": 0.6405211366713047
201
+ }
202
+ }
203
+ ],
204
+ "decoder": {
205
+ "temperature": 0.38,
206
+ "radius": 8,
207
+ "epsilon": 0.001
208
+ },
209
+ "selection_scope": "50 Imagenette plus200 excluded COCO training-validation images; external tests not used to rank weights",
210
+ "model_sha256": "0e4c417375684a044860f8af3ac3a2fb44e1a5729254ca33abda758ea71aea6e"
211
+ }
requirements-onnx.txt ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ onnxruntime>=1.17
2
+ numpy
3
+ pillow
4
+ scikit-image
requirements.txt ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ torch>=2.3
2
+ huggingface_hub>=0.24
3
+ safetensors>=0.4
4
+ scikit-image>=0.22
5
+ scipy>=1.11
6
+ pillow>=10
7
+ numpy
8
+ pyarrow
9
+ pytest
10
+ datasets
11
+ torchvision>=0.18
12
+ tqdm
spatial.py ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Controlled spatial decoding alternatives; no learned parameters."""
2
+ import math
3
+ import torch
4
+ import torch.nn.functional as F
5
+
6
+ def box_mean(x,radius):
7
+ # Border-normalized local windows, also work on images smaller than kernel.
8
+ return F.avg_pool2d(x,2*radius+1,stride=1,padding=radius,count_include_pad=False)
9
+
10
+ def guided_chroma(L,ab,radius=4,epsilon=.001):
11
+ """Scalar luminance-guided local linear filter (He et al., ECCV 2010).
12
+
13
+ L is the network's [-1,1] luminance. Epsilon is in [0,1] luminance units.
14
+ Uses luminance only; reference colors never enter inference.
15
+ """
16
+ if not isinstance(radius,int) or radius<0 or not math.isfinite(epsilon) or epsilon<=0:
17
+ raise ValueError('Invalid guided-filter radius or epsilon')
18
+ if radius==0:return ab
19
+ I=(L.float()+1)/2;p=ab.float()
20
+ mi=box_mean(I,radius);mp=box_mean(p,radius)
21
+ var=(box_mean(I*I,radius)-mi*mi).clamp_min(0)
22
+ cov=box_mean(I*p,radius)-mi*mp
23
+ a=cov/(var+epsilon);b=mp-a*mi
24
+ return box_mean(a,radius)*I+box_mean(b,radius)
25
+
26
+ def spatial_decode(model,logits,L,temperature=.38,pool=1,radius=0,epsilon=.001):
27
+ if pool<1 or not isinstance(pool,int):raise ValueError('pool must be a positive integer')
28
+ if pool>1:
29
+ # Average evidence before annealing, then upsample chroma, not RGB.
30
+ small=F.avg_pool2d(logits,pool,ceil_mode=True,count_include_pad=False)
31
+ ab=model.decode(small,temperature)
32
+ ab=F.interpolate(ab,size=L.shape[-2:],mode='bilinear',align_corners=False)
33
+ else:ab=model.decode(logits,temperature)
34
+ return guided_chroma(L,ab,radius,epsilon)
training/color_statistics.npz ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1d14920885324137169e63f908a413ebf083e8db1e604d2333847fd2094f36a6
3
+ size 3588
training/environment.txt ADDED
@@ -0,0 +1,125 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ anyio==4.15.1
2
+ archspec @ file:///home/conda/feedstock_root/build_artifacts/archspec_1737352602016/work
3
+ asttokens @ file:///home/conda/feedstock_root/build_artifacts/asttokens_1733250440834/work
4
+ astunparse==1.6.3
5
+ attrs @ file:///home/conda/feedstock_root/build_artifacts/attrs_1737819173731/work
6
+ beautifulsoup4 @ file:///home/conda/feedstock_root/build_artifacts/beautifulsoup4_1733230845337/work
7
+ boltons @ file:///home/conda/feedstock_root/build_artifacts/boltons_1733827268945/work
8
+ Brotli @ file:///home/conda/feedstock_root/build_artifacts/brotli-split_1725267488082/work
9
+ certifi @ file:///home/conda/feedstock_root/build_artifacts/certifi_1734380492396/work/certifi
10
+ cffi @ file:///home/conda/feedstock_root/build_artifacts/cffi_1725560564262/work
11
+ chardet @ file:///home/conda/feedstock_root/build_artifacts/chardet_1724954797915/work
12
+ charset-normalizer @ file:///home/conda/feedstock_root/build_artifacts/charset-normalizer_1735929714516/work
13
+ click==8.5.0
14
+ cmake==3.31.4
15
+ colorama @ file:///home/conda/feedstock_root/build_artifacts/colorama_1733218098505/work
16
+ conda @ file:///home/conda/feedstock_root/build_artifacts/conda_1737546925511/work
17
+ conda-build @ file:///home/conda/feedstock_root/build_artifacts/conda-build_1737106456198/work
18
+ conda-libmamba-solver @ file:///home/conda/feedstock_root/build_artifacts/conda-libmamba-solver_1737800978214/work/src
19
+ conda-package-handling @ file:///home/conda/feedstock_root/build_artifacts/conda-package-handling_1736345463896/work
20
+ conda_index @ file:///home/conda/feedstock_root/build_artifacts/conda-index_1718383105992/work
21
+ conda_package_streaming @ file:///home/conda/feedstock_root/build_artifacts/conda-package-streaming_1729004031731/work
22
+ decorator @ file:///home/conda/feedstock_root/build_artifacts/decorator_1733236420667/work
23
+ distro @ file:///home/conda/feedstock_root/build_artifacts/distro_1734729835256/work
24
+ dnspython==2.7.0
25
+ exceptiongroup @ file:///home/conda/feedstock_root/build_artifacts/exceptiongroup_1733208806608/work
26
+ executing @ file:///home/conda/feedstock_root/build_artifacts/executing_1733569351617/work
27
+ expecttest==0.3.0
28
+ filelock @ file:///home/conda/feedstock_root/build_artifacts/filelock_1737517818712/work
29
+ frozendict @ file:///home/conda/feedstock_root/build_artifacts/frozendict_1728841334936/work
30
+ fsspec==2024.12.0
31
+ h11==0.16.0
32
+ h2 @ file:///home/conda/feedstock_root/build_artifacts/h2_1733298745555/work
33
+ hf-xet==1.6.0
34
+ hpack @ file:///home/conda/feedstock_root/build_artifacts/hpack_1733299205993/work
35
+ httpcore==1.0.9
36
+ httpx==0.28.1
37
+ huggingface_hub==1.32.0
38
+ hyperframe @ file:///home/conda/feedstock_root/build_artifacts/hyperframe_1733298771451/work
39
+ hypothesis==6.124.7
40
+ idna @ file:///home/conda/feedstock_root/build_artifacts/idna_1733211830134/work
41
+ ImageIO==2.37.4
42
+ importlib_resources @ file:///home/conda/feedstock_root/build_artifacts/importlib_resources_1736252299705/work
43
+ ipython @ file:///home/conda/feedstock_root/build_artifacts/ipython_1734788142186/work
44
+ jedi @ file:///home/conda/feedstock_root/build_artifacts/jedi_1733300866624/work
45
+ Jinja2 @ file:///home/conda/feedstock_root/build_artifacts/jinja2_1734823942230/work
46
+ jsonpatch @ file:///home/conda/feedstock_root/build_artifacts/jsonpatch_1733814567314/work
47
+ jsonpointer @ file:///home/conda/feedstock_root/build_artifacts/jsonpointer_1725302941992/work
48
+ jsonschema @ file:///home/conda/feedstock_root/build_artifacts/jsonschema_1733472696581/work
49
+ jsonschema-specifications @ file:///tmp/tmpk0f344m9/src
50
+ lazy-loader==0.5
51
+ libarchive-c @ file:///home/conda/feedstock_root/build_artifacts/python-libarchive-c_1725302626023/work
52
+ libmambapy @ file:///home/conda/feedstock_root/build_artifacts/mamba-split_1735806506118/work/libmambapy
53
+ lief @ file:///home/conda/feedstock_root/build_artifacts/lief_1726040283347/work/api/python
54
+ lintrunner==0.12.7
55
+ MarkupSafe @ file:///home/conda/feedstock_root/build_artifacts/markupsafe_1733219680183/work
56
+ matplotlib-inline @ file:///home/conda/feedstock_root/build_artifacts/matplotlib-inline_1733416936468/work
57
+ menuinst @ file:///home/conda/feedstock_root/build_artifacts/menuinst_1731146975675/work
58
+ more-itertools @ file:///home/conda/feedstock_root/build_artifacts/more-itertools_1736883817510/work
59
+ mpmath==1.3.0
60
+ networkx==3.4.2
61
+ ninja==1.11.1.3
62
+ numpy @ file:///home/conda/feedstock_root/build_artifacts/numpy_1737331065203/work/dist/numpy-2.2.2-cp311-cp311-linux_x86_64.whl#sha256=6384aec253ac1c4a8e4c91c88343657bb0a58927963bd1b7fad53f09d3fa122a
63
+ nvidia-cublas-cu12==12.4.5.8
64
+ nvidia-cuda-cupti-cu12==12.4.127
65
+ nvidia-cuda-nvrtc-cu12==12.4.127
66
+ nvidia-cuda-runtime-cu12==12.4.127
67
+ nvidia-cudnn-cu12==9.1.0.70
68
+ nvidia-cufft-cu12==11.2.1.3
69
+ nvidia-curand-cu12==10.3.5.147
70
+ nvidia-cusolver-cu12==11.6.1.9
71
+ nvidia-cusparse-cu12==12.3.1.170
72
+ nvidia-cusparselt-cu12==0.6.2
73
+ nvidia-nccl-cu12==2.21.5
74
+ nvidia-nvjitlink-cu12==12.4.127
75
+ nvidia-nvtx-cu12==12.4.127
76
+ optree==0.14.0
77
+ packaging @ file:///home/conda/feedstock_root/build_artifacts/packaging_1733203243479/work
78
+ parso @ file:///home/conda/feedstock_root/build_artifacts/parso_1733271261340/work
79
+ pexpect @ file:///home/conda/feedstock_root/build_artifacts/pexpect_1733301927746/work
80
+ pickleshare @ file:///home/conda/feedstock_root/build_artifacts/pickleshare_1733327343728/work
81
+ pillow==11.0.0
82
+ pkginfo @ file:///home/conda/feedstock_root/build_artifacts/pkginfo_1733734533957/work
83
+ pkgutil_resolve_name @ file:///home/conda/feedstock_root/build_artifacts/pkgutil-resolve-name_1733344503739/work
84
+ platformdirs @ file:///home/conda/feedstock_root/build_artifacts/platformdirs_1733232627818/work
85
+ pluggy @ file:///home/conda/feedstock_root/build_artifacts/pluggy_1733222765875/work
86
+ prompt_toolkit @ file:///home/conda/feedstock_root/build_artifacts/prompt-toolkit_1737453357274/work
87
+ psutil @ file:///home/conda/feedstock_root/build_artifacts/psutil_1735327328223/work
88
+ ptyprocess @ file:///home/conda/feedstock_root/build_artifacts/ptyprocess_1733302279685/work/dist/ptyprocess-0.7.0-py2.py3-none-any.whl#sha256=92c32ff62b5fd8cf325bec5ab90d7be3d2a8ca8c8a3813ff487a8d2002630d1f
89
+ pure_eval @ file:///home/conda/feedstock_root/build_artifacts/pure_eval_1733569405015/work
90
+ pyarrow==25.0.1
91
+ pycosat @ file:///home/conda/feedstock_root/build_artifacts/pycosat_1732588400443/work
92
+ pycparser @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_pycparser_1733195786/work
93
+ Pygments @ file:///home/conda/feedstock_root/build_artifacts/pygments_1736243443484/work
94
+ PySocks @ file:///home/conda/feedstock_root/build_artifacts/pysocks_1733217236728/work
95
+ python-etcd==0.4.5
96
+ pytz @ file:///home/conda/feedstock_root/build_artifacts/pytz_1733215667876/work
97
+ PyYAML @ file:///home/conda/feedstock_root/build_artifacts/pyyaml_1737454647378/work
98
+ referencing @ file:///home/conda/feedstock_root/build_artifacts/bld/rattler-build_referencing_1737836872/work
99
+ requests @ file:///home/conda/feedstock_root/build_artifacts/requests_1733217035951/work
100
+ rpds-py @ file:///home/conda/feedstock_root/build_artifacts/rpds-py_1733366613949/work
101
+ ruamel.yaml @ file:///home/conda/feedstock_root/build_artifacts/ruamel.yaml_1736248028599/work
102
+ ruamel.yaml.clib @ file:///home/conda/feedstock_root/build_artifacts/ruamel.yaml.clib_1728724459810/work
103
+ safetensors==0.8.0
104
+ scikit-image==0.26.0
105
+ scipy==1.17.1
106
+ six==1.17.0
107
+ sortedcontainers==2.4.0
108
+ soupsieve @ file:///home/conda/feedstock_root/build_artifacts/soupsieve_1693929250441/work
109
+ stack_data @ file:///home/conda/feedstock_root/build_artifacts/stack_data_1733569443808/work
110
+ sympy==1.13.1
111
+ tifffile==2026.3.3
112
+ torch==2.6.0+cu124
113
+ torchaudio==2.6.0+cu124
114
+ torchelastic==0.2.2
115
+ torchvision==0.21.0+cu124
116
+ tqdm @ file:///home/conda/feedstock_root/build_artifacts/tqdm_1735661334605/work
117
+ traitlets @ file:///home/conda/feedstock_root/build_artifacts/traitlets_1733367359838/work
118
+ triton==3.2.0
119
+ truststore @ file:///home/conda/feedstock_root/build_artifacts/truststore_1729762363021/work
120
+ types-dataclasses==0.6.6
121
+ typing_extensions==4.16.0
122
+ urllib3 @ file:///home/conda/feedstock_root/build_artifacts/urllib3_1734859416348/work
123
+ wcwidth @ file:///home/conda/feedstock_root/build_artifacts/wcwidth_1733231326287/work
124
+ zipp @ file:///home/conda/feedstock_root/build_artifacts/zipp_1732827521216/work
125
+ zstandard==0.23.0
training/history.json ADDED
@@ -0,0 +1,101 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [
2
+ {
3
+ "step": 0,
4
+ "validation": {
5
+ "ce": 3.0858358535766603,
6
+ "ab_error": 14.951502853393555,
7
+ "chroma": 13.526052837371827,
8
+ "target_chroma": 14.882324807823636,
9
+ "excess_chroma_edge": 0.4721723803579807,
10
+ "n": 250,
11
+ "chroma_ratio": 0.908866928523236,
12
+ "seconds": 10.445266901000025,
13
+ "ms_per_image": 41.7810676040001,
14
+ "low_chroma_sources": 11,
15
+ "color_sources": {
16
+ "ce": 3.1215308959514028,
17
+ "ab_error": 15.274329947627239,
18
+ "chroma": 13.770119405690597,
19
+ "target_chroma": 15.526453359356486,
20
+ "excess_chroma_edge": 0.45490627698444425,
21
+ "n": 239,
22
+ "chroma_ratio": 0.886881188316745
23
+ }
24
+ }
25
+ },
26
+ {
27
+ "step": 374,
28
+ "loss": 2.872464179992676,
29
+ "validation": {
30
+ "ce": 3.0596936688423155,
31
+ "ab_error": 14.627770241737366,
32
+ "chroma": 13.950701038837433,
33
+ "target_chroma": 14.882324807823636,
34
+ "excess_chroma_edge": 0.42505902045965194,
35
+ "n": 250,
36
+ "chroma_ratio": 0.9374006560791868,
37
+ "seconds": 5.679420730999993,
38
+ "ms_per_image": 22.71768292399997,
39
+ "low_chroma_sources": 11,
40
+ "color_sources": {
41
+ "ce": 3.093559457667203,
42
+ "ab_error": 14.939232912023696,
43
+ "chroma": 14.216437334296096,
44
+ "target_chroma": 15.526453359356486,
45
+ "excess_chroma_edge": 0.408694201371161,
46
+ "n": 239,
47
+ "chroma_ratio": 0.9156268341043285
48
+ }
49
+ }
50
+ },
51
+ {
52
+ "step": 748,
53
+ "loss": 2.7719690799713135,
54
+ "validation": {
55
+ "ce": 3.0443123207092286,
56
+ "ab_error": 14.446234788894653,
57
+ "chroma": 13.278011664867401,
58
+ "target_chroma": 14.882324807823636,
59
+ "excess_chroma_edge": 0.45324007369577884,
60
+ "n": 250,
61
+ "chroma_ratio": 0.8922000988640667,
62
+ "seconds": 5.6056951999999,
63
+ "ms_per_image": 22.4227807999996,
64
+ "low_chroma_sources": 11,
65
+ "color_sources": {
66
+ "ce": 3.077849601601956,
67
+ "ab_error": 14.765882765398864,
68
+ "chroma": 13.530982503831137,
69
+ "target_chroma": 15.526453359356486,
70
+ "excess_chroma_edge": 0.43782363727564094,
71
+ "n": 239,
72
+ "chroma_ratio": 0.8714792870373809
73
+ }
74
+ }
75
+ },
76
+ {
77
+ "step": 1122,
78
+ "loss": 2.405717611312866,
79
+ "validation": {
80
+ "ce": 3.056647213935852,
81
+ "ab_error": 14.650493629455566,
82
+ "chroma": 14.005748982429504,
83
+ "target_chroma": 14.882324807823636,
84
+ "excess_chroma_edge": 0.45751794615387914,
85
+ "n": 250,
86
+ "chroma_ratio": 0.941099536751589,
87
+ "seconds": 5.821004607999953,
88
+ "ms_per_image": 23.28401843199981,
89
+ "low_chroma_sources": 11,
90
+ "color_sources": {
91
+ "ce": 3.088832367414211,
92
+ "ab_error": 14.943533709857254,
93
+ "chroma": 14.253979418567035,
94
+ "target_chroma": 15.526453359356486,
95
+ "excess_chroma_edge": 0.44000523674562886,
96
+ "n": 239,
97
+ "chroma_ratio": 0.918044777429957
98
+ }
99
+ }
100
+ }
101
+ ]
training/run_config.json ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "init_from": "init",
3
+ "revision": null,
4
+ "parquet": "/tmp/colorizer/combined.parquet",
5
+ "manifest": "manifest.json",
6
+ "output": "run",
7
+ "steps": 1122,
8
+ "size": 256,
9
+ "batch_size": 32,
10
+ "lr": 1e-05,
11
+ "context_lr": 1e-05,
12
+ "head_lr": 1e-05,
13
+ "rebalance_lambda": 0.7,
14
+ "consistency": 0.0,
15
+ "statistics": "stats.npz",
16
+ "freeze_bn": true,
17
+ "seed": 123,
18
+ "threads": 4,
19
+ "workers": 4,
20
+ "eval_every": 374,
21
+ "push_to_hub": false,
22
+ "hub_model_id": "User-2468/mini-unet-colorizer",
23
+ "device": "cuda",
24
+ "amp_dtype": "torch.bfloat16",
25
+ "parameters": 3968892,
26
+ "bin_sha256": "4bcdeb9ad2dcf6d460e7157e436a524b6b885ef8b5552a2f0480550db811e173"
27
+ }
training/split_manifest.json ADDED
The diff for this file is too large to render. See raw diff
 
upload_main.py ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Upload a prepared release folder atomically to main, preserving stable."""
2
+ import argparse
3
+ from pathlib import Path
4
+ from huggingface_hub import HfApi
5
+ from model import load_model
6
+
7
+ def main():
8
+ p=argparse.ArgumentParser(__doc__)
9
+ p.add_argument('--folder',default='release')
10
+ p.add_argument('--repo',default='User-2468/mini-unet-colorizer')
11
+ p.add_argument('--expected-main',default=None,help='Optional optimistic-concurrency commit SHA')
12
+ a=p.parse_args()
13
+ folder=Path(a.folder)
14
+ model=load_model(folder)
15
+ assert sum(p.numel() for p in model.parameters())<4_000_000
16
+ api=HfApi()
17
+ refs={x.name:x.target_commit for x in api.list_repo_refs(a.repo).branches}
18
+ stable=refs.get('stable')
19
+ result=api.upload_folder(repo_id=a.repo,revision='main',folder_path=folder,
20
+ parent_commit=a.expected_main or refs['main'],
21
+ commit_message='Update validated colorizer weights, spatial decoder and deployment artifacts',
22
+ ignore_patterns=['__pycache__/**','*.pyc'])
23
+ after={x.name:x.target_commit for x in api.list_repo_refs(a.repo).branches}
24
+ if stable!=after.get('stable'):raise RuntimeError('Stable changed concurrently; inspect repository history')
25
+ print(result.commit_url)
26
+
27
+ if __name__=='__main__':main()