Download README.md from evankuo/Qwen-Image-2.1-MNN: direct link, hf CLI and curl.
- Browser
- Download file 3.5 kB
-
https://huggingface.co/evankuo/Qwen-Image-2.1-MNN/resolve/75469e93ca983b614544075d8c670a7f317a67ef/README.md
- Command line
-
hf download hf://evankuo/Qwen-Image-2.1-MNN@75469e93ca983b614544075d8c670a7f317a67ef/README.md
-
curl -L -o README.md https://huggingface.co/evankuo/Qwen-Image-2.1-MNN/resolve/75469e93ca983b614544075d8c670a7f317a67ef/README.md
license: other
license_name: qwen-research
license_link: LICENSE
base_model:
- Qwen/Qwen-Image-2.1
base_model_relation: quantized
pipeline_tag: text-to-image
library_name: mnn
tags:
- mnn
- android
- opencl
- int4
- qwen-image
Qwen-Image-2.1 · MNN (int4) for Android
Qwen/Qwen-Image-2.1 converted to MNN for on-device text-to-image and image editing on Android, with the OpenCL GPU running the DiT. Any size with sides a multiple of 32 works, from 256×256 up; the app offers 7 aspect ratios at three pixel budgets (~512², ~384², ~320²), e.g. 512×512, 576×448, 672×384, 480×320, 384×288.
Runtime, Android library and demo app: github.com/scsonic/libQwenImage21
| Tested on | Snapdragon 8 Gen 2 (Adreno 740), 16 GB RAM, Android 13 |
|---|---|
| Text to image, 448×576, 20 steps | 451 s total (DiT 19.1 s/step on OpenCL fp16) |
| Image edit, 352×448, 20 steps | 348 s total (DiT 12.6 s/step) |
Files
| Path | What | Size |
|---|---|---|
dit.mnn + .weight |
7B single-stream DiT (32 blocks + norm_out/proj_out). Block linears int4 (block 32), small layers int8 | 4.5 GB |
txt_in.mnn, img_in.mnn |
text (int8) / latent (fp16) input projections | 36 MB |
vae_decoder.mnn |
VAE decoder, 64-ch latent → RGBA, fp16 weights, dynamic size | 0.5 GB |
vae_encoder.mnn |
VAE encoder for image editing, RGBA → normalized 64-ch latent, fp16 | 0.16 GB |
text_encoder/ |
Qwen3-VL-8B-Instruct, MNN int4 (from taobao-mnn/Qwen3-VL-8B-Instruct-MNN). te_config.json runs it text-only; te_vl_config.json adds the vision tower (visual.mnn) for image editing. Both return the last decoder layer before the final norm |
5.4 GB |
Download:
hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21
How it was made
- DiT: taken from the GGUF Q4_K build (leejet/Qwen-Image-2.1-GGUF).
Every Q4_K sub-block of 32 weights (
w = d·sc·q − dmin·m) maps exactly onto MNN's asymmetric int4 with block 32, so the weights are copied without re-quantization (scales stored as fp16). - Text encoder: Qwen-Image-2.1's
text_encoderis byte-identical to Qwen3-VL-8B-Instruct, so the existing MNN export is reused unchanged. - VAE: the residual stream is divided by 256 (exact, power of two) and RMSNorm pre-divides by max|x| so the decoder fits fp16 (it peaks at ~3.5e5 otherwise). The decoded image is unchanged.
- The pipeline caches the text K/V once per prompt (Qwen-Image-2.1's block-causal attention), so each denoising step
only runs the image tokens. The cache is one tensor per layer (
past_kv_0…past_kv_31): a single[32, 2, P, 32, 128]tensor is exactly 1 MiB per prefix token, and an image-edit prefix (P > 1000) would exceed OpenCL's 1 GiB maximum buffer size on an Adreno 740.
2026-09-23:
dit.mnn/dit.mnn.weightwere re-exported for that per-layer K/V cache. Older copies do not load with the current runtime — re-download both files.
Conversion scripts: export/ in the GitHub repo.
License
Derived from Qwen-Image-2.1 and released under the Qwen Research License Agreement (see LICENSE), i.e. for
research / non-commercial use under its terms. The text encoder weights come from Qwen3-VL-8B-Instruct (Apache-2.0).