Ming Image 0.1 Design — W4A8 ConvRot (experimental)
A ready-to-download W4A8 diffusion transformer (DiT) for inclusionAI/Ming-Image-0.1-Design, converted and tested for ComfyUI / Aikimi Forge Neo. Users do not need to convert the model. This is a community quantization, not an official inclusionAI release.
Only the DiT is distributed here. Use the companion W4A8 text encoder and BF16 VAE from Comfy-Org/Ming-Image.
Download and use
Aikimi Forge Neo
In v3.2.2 or later, open Ming Image → 詳細設定 → 本体モデル → W4A8(省メモリ・試験版), then press Ming Imageを準備. Neo downloads the checkpoint and conversion record from this repository, plus the shared encoder/VAE from Comfy-Org, with pinned revisions and SHA-256 verification. The selected W4A8 setup does not require the INT8 DiT or BF16 source.
Alternatively, from the Neo directory:
.\venv\Scripts\python.exe tools\setup_ming_image.py --precision w4a8
ComfyUI
- Download
diffusion_models/ming_image_0.1_design_w4a8_convrot_experimental.safetensorsintoComfyUI/models/diffusion_models/. - Download
text_encoders/ming_image_0.1_ling_mini_2.0_w4a8.safetensorsandvae/ming_image_vae_bf16.safetensorsfrom the companion repository into the matching ComfyUI model directories. - Use a Ming-compatible workflow with the W4A8 diffusion model, the Ling-mini encoder, and the Ming VAE. Tested with ComfyUI
3b4c0b0e457cf0a51cf3038e0a6750d8f96ce251(0.37.0), comfy-kitchen 0.2.35, PyTorch 2.11.0+cu130, Python 3.12.13 on Windows / NVIDIA CUDA.
The JSON alongside the checkpoint records its provenance and integrity. Neo requires both files; it downloads them together. Recommended test settings are 1024×1024, Euler/simple, 12 steps and CFG 1. For transparent output, the tested 2048×2048 case worked better; transparency is prompt- and resolution-dependent.
What was quantized
- Source: Comfy-Org BF16 DiT at revision
53654871e47a5d2daed7b3a986cbf1010ef81c78, originally by inclusionAI. Not re-quantized from INT8. - Native Q/K/V fusion using the pinned ComfyUI mapping before quantization.
- 202 linear layers,
asym_w4a8_int8, group size 16, ConvRot group size 256. Non-quantized layers retain their original precision. - Converter: Comfy-Org/comfy-model-tools, revision
d6797787e6bdb1a1fb0094d588a26f8e71a1c757. - Reproduction helper: tools/quantize_ming_image.py. This is optional; downloading this checkpoint is sufficient.
- Output size: 3,488,163,008 bytes (3.49 GB).
- Output SHA-256:
66be75b57dfb464a905f7f1359a63303890e08ff8864d9fb8470e2513030d515. - BF16 source SHA-256:
8781c6fc679f1812ad318cc7fd84c4d05a23b3571b9a7efeff73a92054ef4659.
Attention metadata is retained to match the distributed INT8 checkpoint. The tested Ming runtime reports those keys as unused for both variants; their presence is not evidence that an INT8 attention kernel ran. The W4A8 linear layers loaded and generated images successfully.
RTX 3090 comparison
Measured 2026-09-29, RTX 3090 24GiB / system RAM 64GB, same prompts, seeds, 12 steps and CFG 1. Text encoder is W4A8 and VAE BF16 in both columns.
| Metric | INT8 DiT | W4A8 DiT |
|---|---|---|
| DiT file | 6.18 GB | 3.49 GB |
| 1024×1024 posters, warm model, two seeds | 7.4–7.6 s | 8.5–8.7 s |
| 1024 poster sampled GPU maximum | 20.2 GiB | 17.9 GiB |
| 2048×2048 transparent leaf, warm model | 53.2 s | 57.8 s |
| 2048 leaf sampled GPU maximum | 21.2 GiB | 20.9 GiB |
W4A8 reduced memory use noticeably in these 1024px runs, but the 2K peak changed little and generation became slower. Times exclude runtime startup and pre-job integrity checks. GPU figures include other applications and are sampled maxima, not guaranteed peaks. These are a few individual measurements, not a broad benchmark or minimum-VRAM guarantee.
Both poster variants rendered the headline/subtitle legibly, but the W4A8 example below added unwanted small text at the lower right. Composition and fine detail can change. Both tested 2048px leaf outputs contained real alpha transparency. There is no guarantee of matching INT8 quality or accurate lettering.
Full experiment record and conditions.
License and credits
MIT; the original inclusionAI copyright and license are included in LICENSE. Original model: inclusionAI. ComfyUI packaging/BF16 source: Comfy-Org / Kijai. Quantization tooling/runtime: Comfy-Org, comfy-kitchen and ComfyUI contributors. This repository supplies an experimental derived checkpoint and its conversion record.
Model tree for Aikimi/Ming-Image-0.1-Design-W4A8
Base model
inclusionAI/Ming-Image-0.1-Design