--- license: mit base_model: inclusionAI/Ming-Image-0.1-Design library_name: mflux pipeline_tag: text-to-image tags: - mlx - mflux - apple-silicon - text-to-image - graphic-design - rgba --- # Ming-Image-0.1-Design — MLX (8-bit DiT, 8-bit text encoder) An MLX conversion of [inclusionAI/Ming-Image-0.1-Design](https://huggingface.co/inclusionAI/Ming-Image-0.1-Design) for [mflux](https://github.com/filipstrand/mflux) — 25.2 GB in total. Ming-Image is a design-focused text-to-image model (posters, cards, UI, typography) that outputs **RGBA** images. | Component | Precision | Size | |---|---|---| | Ling-mini-2.0 MoE text encoder (`mllm/`) | 8-bit (routers, norms, query tokens full precision) | 17.0 GB | | Qwen2 connector (`connector/`) | 8-bit | 1.4 GB | | Projection heads (`mlp/`) | bf16 | 0.06 GB | | DiT (`transformer/`) | 8-bit | 6.5 GB | | RGBA VAE (`vae/`) | bf16 | 0.25 GB | The checkpoint's Qwen2.5 ViT, `lm_head` and audio router are not used for text-to-image and are not included. ## Usage Requires mflux with Ming-Image support (the `ming-image` model family; until it is merged upstream, install mflux from the branch that adds it). ```sh mflux-generate-ming \ --model path/to/this/repo \ --prompt "A modern tech conference poster titled 'MLX SUMMIT 2026' ..." \ --width 1024 --height 1024 --seed 42 ``` Defaults follow the official pipeline: 12 steps, guidance 1.0, flow-match Euler with the checkpoint's static shift of 6. PNGs keep the alpha channel (`--flatten-alpha` for RGB). ## Performance The DiT stage is identical to the 5-bit-text-encoder variant: on a base M4 Mac mini (24 GB) 512² takes ~1 min (9.8 GB peak) and 1024² 4–4.5 min (14.5 GB peak). The text encoder is freed after encoding the prompt. ## Fidelity The MLX implementation was checked stage by stage against tensors dumped from the official PyTorch pipeline: DiT, VAE and connector at cosine similarity ≥ 0.9999, and the text encoder within the official model's own eager-vs-flash-attention spread (this includes reproducing the official pipeline's autocast-bf16 MoE routing exactly). From the same initial noise it reproduces the official images' layout, typography and text. The text-encoder stage needs roughly 18.5 GB (estimated from the weight sizes), so use a Mac with 32 GB or more. The DiT is deliberately kept at 8 bits: at 4 bits large headline lettering degrades. ## License MIT, as the original model. Credit to inclusionAI for [Ming-Image](https://github.com/inclusionAI/Ming-Image).