--- license: other license_name: qwen-research license_link: LICENSE base_model: - Qwen/Qwen-Image-2.1 base_model_relation: quantized pipeline_tag: text-to-image library_name: mnn tags: - mnn - android - opencl - int4 - qwen-image --- # Qwen-Image-2.1 · MNN (int4) for Android [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) converted to [MNN](https://github.com/alibaba/MNN) for on-device text-to-image at 512×512 on Android (OpenCL GPU for the DiT). Runtime, Android library and demo app: **[github.com/scsonic/libQwenImage21](https://github.com/scsonic/libQwenImage21)** | Tested on | Snapdragon 8 Gen 2 (Adreno 740), 16 GB RAM, Android 13 | |---|---| | 512×512, 20 steps | ~490 s total (DiT ~19.7 s/step on OpenCL fp16) | ## Files | Path | What | Size | |---|---|---| | `dit.mnn` + `.weight` | 7B single-stream DiT (32 blocks + norm_out/proj_out). Block linears int4 (block 32), small layers int8 | 4.5 GB | | `txt_in.mnn`, `img_in.mnn` | text / latent input projections (int8) | 36 MB | | `vae_decoder.mnn` | VAE decoder, 64-ch latent → RGBA, fp16 weights | 0.5 GB | | `text_encoder/` | Qwen3-VL-8B-Instruct, MNN int4 (from [taobao-mnn/Qwen3-VL-8B-Instruct-MNN](https://huggingface.co/taobao-mnn/Qwen3-VL-8B-Instruct-MNN)); `te_config.json` runs it text-only and returns the last decoder layer before the final norm | 5.1 GB | Download: ```bash hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21 ``` ## How it was made - **DiT**: taken from the GGUF **Q4_K** build ([leejet/Qwen-Image-2.1-GGUF](https://huggingface.co/leejet/Qwen-Image-2.1-GGUF)). Every Q4_K sub-block of 32 weights (`w = d·sc·q − dmin·m`) maps exactly onto MNN's asymmetric int4 with block 32, so the weights are copied without re-quantization (scales stored as fp16). - **Text encoder**: Qwen-Image-2.1's `text_encoder` is byte-identical to Qwen3-VL-8B-Instruct, so the existing MNN export is reused unchanged. - **VAE**: the residual stream is divided by 256 (exact, power of two) and RMSNorm pre-divides by max|x| so the decoder fits fp16 (it peaks at ~3.5e5 otherwise). The decoded image is unchanged. - The pipeline caches the text K/V once per prompt (Qwen-Image-2.1's block-causal attention), so each denoising step only runs the 1024 image tokens. Conversion scripts: `export/` in the GitHub repo. ## License Derived from Qwen-Image-2.1 and released under the **Qwen Research License Agreement** (see `LICENSE`), i.e. for research / non-commercial use under its terms. The text encoder weights come from Qwen3-VL-8B-Instruct (Apache-2.0).