Qwen-Image-2.1 Q4_K_M for Apple Silicon (MLX)

Generate images locally on a Mac with Apple Silicon using Qwen-Image-2.1. This experimental package includes approximately 9.52 GB of weights, a ready-to-run Python inference script, setup instructions and example outputs. Tested on a MacBook Air M3 with 16 GB unified memory at 512 × 512.

Use the included runner: these packed weights require custom MLX kernels and are not a drop-in model for the standard MLX or MFLUX loader.

An Apple Silicon execution port of Unsloth's Qwen-Image-2.1 Q4_K_M image transformer and Qwen3-VL-8B UD-Q4_K_XL text encoder, using MFLUX architecture classes and mlx-kquant Metal kernels.

The packed tensor payloads are copied byte-for-byte from the pinned Unsloth GGUF files. There is no retraining and no second quantization pass. This is a custom packed format: installing MLX alone, or loading these files with the standard MFLUX affine-quantization loader, is insufficient. Use the included runner.

What this release contributes

  • Mixed GGUF codecs inside an MLX image-generation pipeline.
  • Exact source-to-MLX tensor mapping, including fused gate/up matrices.
  • A layer-streamed text encoder and resident image transformer.
  • Lossless, sharded packed safetensors plus source hashes and a verifier.
  • A memory-guarded runner, independent codec/block checks and reproducible examples.

The underlying models belong to their upstream authors; Unsloth created the source quantizations, MFLUX implements the architecture and mlx-kquant supplies the packed kernels. See NOTICE.md for attribution and pinned revisions.

Download

The complete model folder includes the runner and verification source, so GitHub access is not required to use it. After creating a Python3.11 environment:

uv venv .venv --python 3.11
uv pip install --python .venv/bin/python huggingface_hub==1.33.0
.venv/bin/python -c "from huggingface_hub import snapshot_download; snapshot_download('Nurymanau/Qwen-Image-2.1-Q4_K_M-MLX', local_dir='packed-model')"
uv pip install --python .venv/bin/python -r packed-model/requirements.lock
uv pip install --python .venv/bin/python --no-deps 'git+https://github.com/mflux-community/mflux.git@9ca480fd87fce5623e90878766c751d9140c4aa8'

The download is about 9.52 GB. Run the commands below from the same directory. Optional full-file integrity check on macOS: (cd packed-model && shasum -a 256 -c SHA256SUMS).

Usage

See USAGE.md for environment installation, generation, direct GGUF inference, artifact reconstruction and verification. Tested runtime: Python 3.11, MLX 0.32.1, mlx-kquant 0.4.13, pinned MFLUX 0.20.0 architecture, macOS 26.6.2, MacBook Air M3 with 16 GiB unified memory.

.venv/bin/python packed-model/benchmark.py --output runs/example --timeout 1800 -- \
  .venv/bin/python packed-model/generate.py --encoder packed-model --transformer packed-model \
  --tokenizer packed-model/tokenizer.json --vae packed-model/vae.safetensors \
  --output runs/example/artifacts --size 512 --steps 40 --vae-tiling

Evidence and limits

See RESULTS.md for actual measurements, failed attempts and verification scope. A completed image demonstrates one configuration on one Mac; it is not a universal quality, speed or zero-swap guarantee. No claim of first implementation or superiority over other backends is made.

Only text-to-image is qualified here. Image editing, LoRA, other quantizations, other resolutions and other hardware remain unqualified. The original VAE and tokenizer are included as separate unchanged assets. The unused encoder logits head is explicitly excluded; the exported inference tensors total 8,831,774,720 bytes before VAE/tokenizer/metadata.

License

The image model is governed by the Qwen RESEARCH LICENSE AGREEMENT, including its non-commercial research/evaluation restriction. Commercial use requires a separate upstream license. The text encoder retains Apache-2.0 terms; adapter code and dependencies have their own licenses. Do not treat the adapter's MIT license as permission to commercially use the image model.

Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved.

Packed U8 tensors store compressed bytes. Automatic tensor-element counts on a hosting site should not be interpreted as the number of logical model parameters.

Additional evidence

Three-prompt comparison with original GGUF · Measured encoder optimization. Complete outputs and criterion decisions are included.

Product,40 steps

Exact text,16 steps

Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Nurymanau/Qwen-Image-2.1-Q4_K_M-MLX

Finetuned
(54)
this model