Qwen-Image-2.1 · MLX 8-bit (mflux)

Qwen/Qwen-Image-2.1 quantized to 8-bit for Apple Silicon with mflux (mflux-save -q 8, commit 8c00dab2).

Component Precision Size
Transformer (7B single-stream DiT) 8-bit (MLX affine, group 64) 7.1 GiB
Text encoder (Qwen3-VL-8B) bf16 (mflux keeps it unquantized) 14 GiB
VAE (64-channel) 8-bit linear layers, conv layers bf16 1.2 GiB
Total ~23 GB

Usage

Requires an mflux build that includes Qwen-Image-2.1 support (mflux main at or after 8c00dab2).

mflux-generate-qwen-2.1 \
  --model JoyFusionAI/Qwen-Image-2.1-MLX-8bit \
  --base-model qwen-image-2.1 \
  --prompt "A red fox in a snowy forest holding a wooden sign that says \"Hello\"" \
  --steps 40 --low-ram

--low-ram releases the text encoder after prompt encoding and tiles the VAE decode. Measured on an M1 Max at 768×768: peak memory ~10 GB (vs. 33–36 GB without it), ~6.7 s/step, no visible quality change.

Defaults recommended by the model authors: 40 steps, no guidance. Image sizes must be divisible by 16.

License

These weights are a derivative of Qwen-Image-2.1 and are distributed under the Qwen Research License Agreement. Non-commercial research and evaluation use only; commercial use requires a separate license from the Qwen team. See the original license.

Qwen-Image-2.1 is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) Alibaba Cloud. All Rights Reserved.

Downloads last month
3,186
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JoyFusionAI/Qwen-Image-2.1-MLX-8bit

Quantized
(85)
this model