--- license: apache-2.0 base_model: Qwen/Qwen-Image-2.1 base_model_relation: quantized library_name: mlx-serve tags: - mlx - mlx-serve - quantized - text-to-image - image-to-image pipeline_tag: text-to-image --- # Qwen-Image-2.1 MLX-Serve 4-bit 4-bit pack of [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) for [mlx-serve](https://github.com/ddalcu/mlx-serve): 11.2 GB, for 16 GB Macs. ![sample](https://raw.githubusercontent.com/ddalcu/mlx-serve/main/website/screenshots/qwen-image-2.1-4bit-512.jpg) ## What is in it The checkpoint's own diffusers layout and key names, with the DiT block linears and the text-encoder layer linears affine-quantized to 4-bit (group 64). Kept dense: the VAE (f32), `embed_tokens`, norms, and the DiT's small or shared linears. Kept: the Qwen3-VL vision tower, for instruction editing. Dropped: `lm_head` and the VAE's per-frame `time_conv`s. Built by `tests/convert_qwen_image21_weights.py --preset 16gb`. ## Measured (M1 Pro, 32 GB) | Pack | Size | Steps | Wall clock incl. load | Peak memory | |---|---|---|---|---| | 8-bit | 1024x1024 | 40 | 985 s (~23 s/step) | 12.95 GB | | 4-bit | 1024x1024 | 3 | 87 s | 9.55 GB | | 4-bit | 512x512 | 20 | 118 s | - | On a Mac the full set would crowd, mlx-serve loads the text encoder per request and frees it before the denoise, so the resident set is the DiT and VAE. ## Run it ```sh brew tap ddalcu/mlx-serve https://github.com/ddalcu/mlx-serve brew install mlx-serve mlx-serve pull ddalcu/Qwen-Image-2.1-MLX-Serve-4bit mlx-serve serve curl localhost:11234/v1/images/generations -H 'Content-Type: application/json' \ -d '{"model":"ddalcu/Qwen-Image-2.1-MLX-Serve-4bit","prompt":"a red fox in fresh snow","size":"1024x1024"}' ``` 40 steps when `steps` is omitted. `guidance_scale` above 1 with a `negative_prompt` runs real CFG (two forwards per step). `image` + `strength` does image-to-image. `"mode":"edit"` with an `image` (plus up to 9 `ref_images`) edits it from the prompt. Apache-2.0, same as the base model.