Mini Photo Colorizer — v3.0

A compact automatic photo colouriser with 3,994,676 learned parameters in total, including the complete deployed MobileNetV3 encoder. It predicts plausible colours from grayscale photographs. No teacher, critic, external semantic model, retrieval service or ensemble is needed at inference.

Current status — 29 September 2026

Try the ZeroGPU Space.

Root weights and ONNX are the reviewed v3.0 production release. The completed 3,995,832-parameter four-palette continuation was not promoted after final checkpoint review: it lowered patch measures modestly but reduced colour coverage and left historical photos muted. Its selected checkpoint and all training artifacts remain at experiments/final-20260929. See the experiment decision and configuration.

The updated Space offers Standard, Detailed and Maximum input detail, full-resolution PNG output, and cached strength adjustments. Higher input detail can help some photos but can also change their colours. The default stays at256 pixels. The Space can load compatible multi-palette checkpoints for experiments without downloading arbitrary Python code; its default is v3.

Historical research documents have moved to reports/history/. See the report index. The stable branch and existing release weights are preserved.

Release contents

  • model.safetensors and config.json: the selected checkpoint.
  • semantic_model.py: the complete architecture and strict checkpoint loader.
  • inference.py: Python API and command-line image colourisation.
  • colorizer.onnx: equivalent fixed 256×256 network graph, with float Lab ab output.
  • app.py and requirements-space.txt: the tested Gradio / ZeroGPU application.
  • RELEASE_REPORT.md, QA.json and SHA256SUMS.json: selection evidence, runtime checks and file hashes.

Scratch Base93 export

The complete production v3 model is also available as a Base93 text export, including its MobileNetV3 encoder. weights_base93.txt is 4,946,819 bytes, or 4,949,484 bytes as a JSON string including escaping. It uses one character for 87.94% of learned parameters and two for sensitive weights. Python and JavaScript decoders, format documentation and output-preservation evaluation are included. This is lossy storage quantisation; the FP32 checkpoint remains the standard runtime release.

Use

Download this repository, then:

pip install -r requirements.txt
python inference.py input.jpg output.png --model . --device cpu

For CUDA, use --device cuda. Python API:

from PIL import Image
from inference import load_colorizer, colorize

model = load_colorizer('.', 'cpu')
output = colorize(model, Image.open('input.jpg'))
output.save('output.png')

Default processing uses a 256-pixel maximum network side and retains aspect ratio. The output keeps the input resolution, orientation and alpha, up to 12 megapixels. Original Lab lightness is retained; colour is upsampled with a lightness-guided local linear model and compressed into the sRGB gamut. Quantisation can cause small lightness differences.

saturation=1.0 is the default. Values between 0 and 1.5 are supported. Smoothing radius defaults to 8 at network resolution; 4 and 16 are available for gentler or stronger smoothing.

Architecture and training

MobileNetV3 Large feature encoder, 128-channel feature pyramid, 16 colour queries and two attention blocks. The decoder produces a shared palette and spatial assignment masks with a bounded local residual. The ImageNet classifier is discarded, and every remaining learned parameter is included in the count above. There are 5,324 parameters of headroom below the strict four-million limit.

Training uses 16,230 filtered photographs from pinned Imagenette and COCO parquet files, with ImageNet pretrained encoder features and DDColor artistic targets. The final selection and any fine-tuning are documented in RELEASE_REPORT.md. Data revisions, training scripts and historical experiments are retained under experiments/.

Teacher source and checkpoint references:

  • DDColor, artistic checkpoint revision aa10f72fffc89a6658e37b48556050b4d9a26f63.
  • TorchVision MobileNetV3 Large, ImageNet V2 encoder initialization.
  • PalGAN informed palette and realism experiments; this implementation is not a reproduction.

Intended use and limits

Intended for adding plausible colour to ordinary photographs and consumer photo applications. A grayscale image can admit several equally plausible colours. Outputs must not be presented as recovered historical fact.

Known limitations include muted or warm/sepia colour choices, incorrect clothing and object hues, residual colour bleeding around small objects, and poor results on unusual surfaces or scenes. Severe scan damage, astronomical images and illustrations are outside the principal evaluation domain. The model does not repair scratches or reconstruct lost detail. Surface consistency is not the same as assigning a uniform colour to an entire patterned object.

Metrics against an original colour image are auxiliary because that original may not be inferable. The report includes visual review and colour-retention diagnostics. Upstream pretraining overlap is possible; the evaluation is not a guarantee of unseen content or universal quality. No claim is made that further improvements below four million parameters are impossible.

Migration from the old U-Net

This release changes the architecture. The old U-Net class and colour-bin decoder cannot load these weights. Use semantic_model.py and inference.py together with the new checkpoint. The network directly returns Lab ab; do not apply the old bin softmax or temperature decoder. Prior main commits and the user-managed stable branch remain available. Old root-level training and ONNX helpers are archived under legacy/v2/ to avoid silently mixing architectures.

ONNX

The ONNX graph accepts float32 L shaped [1,1,256,256], normalized as L_lab / 50 - 1, and returns float32 Lab ab shaped [1,2,256,256]. It contains only the compact network. Image loading, resizing, guided upsampling and gamut compression are handled by the Python pipeline, not embedded in the graph. See colorizer.json for its contract. The fixed-size export is numerically compared with PyTorch before release.

Downloads last month
759
Safetensors
Model size
4.02M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using User-2468/mini-unet-colorizer 1

Papers for User-2468/mini-unet-colorizer