|
Download README.md from evankuo/Qwen-Image-2.1-MNN: direct link, hf CLI and curl.
- Browser
- Download file 2.58 kB
-
https://huggingface.co/evankuo/Qwen-Image-2.1-MNN/resolve/c546544e0e66ef79d0f026a3669e386bad5f8262/README.md
- Command line
-
hf download hf://evankuo/Qwen-Image-2.1-MNN@c546544e0e66ef79d0f026a3669e386bad5f8262/README.md
-
curl -L -o README.md https://huggingface.co/evankuo/Qwen-Image-2.1-MNN/resolve/c546544e0e66ef79d0f026a3669e386bad5f8262/README.md
2.58 kB
metadata
license: other
license_name: qwen-research
license_link: LICENSE
base_model:
- Qwen/Qwen-Image-2.1
base_model_relation: quantized
pipeline_tag: text-to-image
library_name: mnn
tags:
- mnn
- android
- opencl
- int4
- qwen-image
Qwen-Image-2.1 · MNN (int4) for Android
Qwen/Qwen-Image-2.1 converted to MNN for on-device text-to-image at 512×512 on Android (OpenCL GPU for the DiT).
Runtime, Android library and demo app: github.com/scsonic/libQwenImage21
| Tested on | Snapdragon 8 Gen 2 (Adreno 740), 16 GB RAM, Android 13 |
|---|---|
| 512×512, 20 steps | ~490 s total (DiT ~19.7 s/step on OpenCL fp16) |
Files
| Path | What | Size |
|---|---|---|
dit.mnn + .weight |
7B single-stream DiT (32 blocks + norm_out/proj_out). Block linears int4 (block 32), small layers int8 | 4.5 GB |
txt_in.mnn, img_in.mnn |
text / latent input projections (int8) | 36 MB |
vae_decoder.mnn |
VAE decoder, 64-ch latent → RGBA, fp16 weights | 0.5 GB |
text_encoder/ |
Qwen3-VL-8B-Instruct, MNN int4 (from taobao-mnn/Qwen3-VL-8B-Instruct-MNN); te_config.json runs it text-only and returns the last decoder layer before the final norm |
5.1 GB |
Download:
hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21
How it was made
- DiT: taken from the GGUF Q4_K build (leejet/Qwen-Image-2.1-GGUF).
Every Q4_K sub-block of 32 weights (
w = d·sc·q − dmin·m) maps exactly onto MNN's asymmetric int4 with block 32, so the weights are copied without re-quantization (scales stored as fp16). - Text encoder: Qwen-Image-2.1's
text_encoderis byte-identical to Qwen3-VL-8B-Instruct, so the existing MNN export is reused unchanged. - VAE: the residual stream is divided by 256 (exact, power of two) and RMSNorm pre-divides by max|x| so the decoder fits fp16 (it peaks at ~3.5e5 otherwise). The decoded image is unchanged.
- The pipeline caches the text K/V once per prompt (Qwen-Image-2.1's block-causal attention), so each denoising step only runs the 1024 image tokens.
Conversion scripts: export/ in the GitHub repo.
License
Derived from Qwen-Image-2.1 and released under the Qwen Research License Agreement (see LICENSE), i.e. for
research / non-commercial use under its terms. The text encoder weights come from Qwen3-VL-8B-Instruct (Apache-2.0).