--- license: apache-2.0 pipeline_tag: image-to-image tags: - colorization - image-colorization - pytorch - onnx - mobilenet-v3 - small-model library_name: pytorch --- # Mini Photo Colorizer — v3.0 A compact automatic photo colouriser with **3,994,676 learned parameters in total**, including the complete deployed MobileNetV3 encoder. It predicts plausible colours from grayscale photographs. No teacher, critic, external semantic model, retrieval service or ensemble is needed at inference. ## Current status — 29 September 2026 [Try the ZeroGPU Space](https://huggingface.co/spaces/User-2468/mini-unet-colorizer). Root weights and ONNX are the reviewed **v3.0 production release**. The completed 3,995,832-parameter four-palette continuation was **not promoted** after [final checkpoint review](reports/FINAL_CHECKPOINT_REVIEW.md): it lowered patch measures modestly but reduced colour coverage and left historical photos muted. Its selected checkpoint and all training artifacts remain at [experiments/final-20260929](experiments/final-20260929). See the [experiment decision](reports/FINAL_TRAINING_REVIEW.md) and [configuration](training/final_config.json). The updated Space offers Standard, Detailed and Maximum input detail, full-resolution PNG output, and cached strength adjustments. Higher input detail can help some photos but can also change their colours. The default stays at256 pixels. The Space can load compatible multi-palette checkpoints for experiments without downloading arbitrary Python code; its default is v3. Historical research documents have moved to `reports/history/`. See [the report index](reports/README.md). The `stable` branch and existing release weights are preserved. ## Release contents - `model.safetensors` and `config.json`: the selected checkpoint. - `semantic_model.py`: the complete architecture and strict checkpoint loader. - `inference.py`: Python API and command-line image colourisation. - `colorizer.onnx`: equivalent fixed 256×256 network graph, with float Lab `ab` output. - `app.py` and `requirements-space.txt`: the tested Gradio / ZeroGPU application. - `RELEASE_REPORT.md`, `QA.json` and `SHA256SUMS.json`: selection evidence, runtime checks and file hashes. ## Scratch Base93 export The complete production v3 model is also available as a [Base93 text export](quantized/base93-v3/README.md), including its MobileNetV3 encoder. [weights_base93.txt](quantized/base93-v3/weights_base93.txt) is 4,946,819 bytes, or **4,949,484 bytes as a JSON string including escaping**. It uses one character for 87.94% of learned parameters and two for sensitive weights. Python and JavaScript decoders, format documentation and output-preservation evaluation are included. This is lossy storage quantisation; the FP32 checkpoint remains the standard runtime release. ## Use Download this repository, then: ```bash pip install -r requirements.txt python inference.py input.jpg output.png --model . --device cpu ``` For CUDA, use `--device cuda`. Python API: ```python from PIL import Image from inference import load_colorizer, colorize model = load_colorizer('.', 'cpu') output = colorize(model, Image.open('input.jpg')) output.save('output.png') ``` Default processing uses a 256-pixel maximum network side and retains aspect ratio. The output keeps the input resolution, orientation and alpha, up to 12 megapixels. Original Lab lightness is retained; colour is upsampled with a lightness-guided local linear model and compressed into the sRGB gamut. Quantisation can cause small lightness differences. `saturation=1.0` is the default. Values between 0 and 1.5 are supported. Smoothing radius defaults to 8 at network resolution; 4 and 16 are available for gentler or stronger smoothing. ## Architecture and training MobileNetV3 Large feature encoder, 128-channel feature pyramid, 16 colour queries and two attention blocks. The decoder produces a shared palette and spatial assignment masks with a bounded local residual. The ImageNet classifier is discarded, and every remaining learned parameter is included in the count above. There are 5,324 parameters of headroom below the strict four-million limit. Training uses 16,230 filtered photographs from pinned Imagenette and COCO parquet files, with ImageNet pretrained encoder features and DDColor artistic targets. The final selection and any fine-tuning are documented in `RELEASE_REPORT.md`. Data revisions, training scripts and historical experiments are retained under `experiments/`. Teacher source and checkpoint references: - [DDColor](https://arxiv.org/abs/2212.11613), artistic checkpoint revision `aa10f72fffc89a6658e37b48556050b4d9a26f63`. - [TorchVision MobileNetV3 Large](https://docs.pytorch.org/vision/0.21/models/generated/torchvision.models.mobilenet_v3_large.html), ImageNet V2 encoder initialization. - [PalGAN](https://arxiv.org/abs/2210.11204) informed palette and realism experiments; this implementation is not a reproduction. ## Intended use and limits Intended for adding plausible colour to ordinary photographs and consumer photo applications. A grayscale image can admit several equally plausible colours. Outputs must not be presented as recovered historical fact. Known limitations include muted or warm/sepia colour choices, incorrect clothing and object hues, residual colour bleeding around small objects, and poor results on unusual surfaces or scenes. Severe scan damage, astronomical images and illustrations are outside the principal evaluation domain. The model does not repair scratches or reconstruct lost detail. Surface consistency is not the same as assigning a uniform colour to an entire patterned object. Metrics against an original colour image are auxiliary because that original may not be inferable. The report includes visual review and colour-retention diagnostics. Upstream pretraining overlap is possible; the evaluation is not a guarantee of unseen content or universal quality. No claim is made that further improvements below four million parameters are impossible. ## Migration from the old U-Net This release changes the architecture. The old U-Net class and colour-bin decoder cannot load these weights. Use `semantic_model.py` and `inference.py` together with the new checkpoint. The network directly returns Lab `ab`; do not apply the old bin softmax or temperature decoder. Prior main commits and the user-managed `stable` branch remain available. Old root-level training and ONNX helpers are archived under `legacy/v2/` to avoid silently mixing architectures. ## ONNX The ONNX graph accepts float32 `L` shaped `[1,1,256,256]`, normalized as `L_lab / 50 - 1`, and returns float32 Lab `ab` shaped `[1,2,256,256]`. It contains only the compact network. Image loading, resizing, guided upsampling and gamut compression are handled by the Python pipeline, not embedded in the graph. See `colorizer.json` for its contract. The fixed-size export is numerically compared with PyTorch before release.