|
Download README.md from evankuo/Qwen-Image-2.1-MNN: direct link, hf CLI and curl.
- Browser
- Download file 3.5 kB
-
https://huggingface.co/evankuo/Qwen-Image-2.1-MNN/resolve/75469e93ca983b614544075d8c670a7f317a67ef/README.md
- Command line
-
hf download hf://evankuo/Qwen-Image-2.1-MNN@75469e93ca983b614544075d8c670a7f317a67ef/README.md
-
curl -L -o README.md https://huggingface.co/evankuo/Qwen-Image-2.1-MNN/resolve/75469e93ca983b614544075d8c670a7f317a67ef/README.md
3.5 kB
| license: other | |
| license_name: qwen-research | |
| license_link: LICENSE | |
| base_model: | |
| - Qwen/Qwen-Image-2.1 | |
| base_model_relation: quantized | |
| pipeline_tag: text-to-image | |
| library_name: mnn | |
| tags: | |
| - mnn | |
| - android | |
| - opencl | |
| - int4 | |
| - qwen-image | |
| # Qwen-Image-2.1 · MNN (int4) for Android | |
| [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) converted to [MNN](https://github.com/alibaba/MNN) | |
| for on-device **text-to-image and image editing** on Android, with the OpenCL GPU running the DiT. Any size with sides | |
| a multiple of 32 works, from 256×256 up; the app offers 7 aspect ratios at three pixel budgets (~512², ~384², ~320²), | |
| e.g. 512×512, 576×448, 672×384, 480×320, 384×288. | |
| Runtime, Android library and demo app: **[github.com/scsonic/libQwenImage21](https://github.com/scsonic/libQwenImage21)** | |
| | Tested on | Snapdragon 8 Gen 2 (Adreno 740), 16 GB RAM, Android 13 | | |
| |---|---| | |
| | Text to image, 448×576, 20 steps | 451 s total (DiT 19.1 s/step on OpenCL fp16) | | |
| | Image edit, 352×448, 20 steps | 348 s total (DiT 12.6 s/step) | | |
| ## Files | |
| | Path | What | Size | | |
| |---|---|---| | |
| | `dit.mnn` + `.weight` | 7B single-stream DiT (32 blocks + norm_out/proj_out). Block linears int4 (block 32), small layers int8 | 4.5 GB | | |
| | `txt_in.mnn`, `img_in.mnn` | text (int8) / latent (fp16) input projections | 36 MB | | |
| | `vae_decoder.mnn` | VAE decoder, 64-ch latent → RGBA, fp16 weights, dynamic size | 0.5 GB | | |
| | `vae_encoder.mnn` | VAE encoder for image editing, RGBA → normalized 64-ch latent, fp16 | 0.16 GB | | |
| | `text_encoder/` | Qwen3-VL-8B-Instruct, MNN int4 (from [taobao-mnn/Qwen3-VL-8B-Instruct-MNN](https://huggingface.co/taobao-mnn/Qwen3-VL-8B-Instruct-MNN)). `te_config.json` runs it text-only; `te_vl_config.json` adds the vision tower (`visual.mnn`) for image editing. Both return the last decoder layer before the final norm | 5.4 GB | | |
| Download: | |
| ```bash | |
| hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21 | |
| ``` | |
| ## How it was made | |
| - **DiT**: taken from the GGUF **Q4_K** build ([leejet/Qwen-Image-2.1-GGUF](https://huggingface.co/leejet/Qwen-Image-2.1-GGUF)). | |
| Every Q4_K sub-block of 32 weights (`w = d·sc·q − dmin·m`) maps exactly onto MNN's asymmetric int4 with block 32, | |
| so the weights are copied without re-quantization (scales stored as fp16). | |
| - **Text encoder**: Qwen-Image-2.1's `text_encoder` is byte-identical to Qwen3-VL-8B-Instruct, so the existing MNN | |
| export is reused unchanged. | |
| - **VAE**: the residual stream is divided by 256 (exact, power of two) and RMSNorm pre-divides by max|x| so the decoder | |
| fits fp16 (it peaks at ~3.5e5 otherwise). The decoded image is unchanged. | |
| - The pipeline caches the text K/V once per prompt (Qwen-Image-2.1's block-causal attention), so each denoising step | |
| only runs the image tokens. The cache is one tensor per layer (`past_kv_0`…`past_kv_31`): a single | |
| `[32, 2, P, 32, 128]` tensor is exactly 1 MiB per prefix token, and an image-edit prefix (P > 1000) would exceed | |
| OpenCL's 1 GiB maximum buffer size on an Adreno 740. | |
| > **2026-09-23:** `dit.mnn` / `dit.mnn.weight` were re-exported for that per-layer K/V cache. Older copies do not load | |
| > with the current runtime — re-download both files. | |
| Conversion scripts: `export/` in the GitHub repo. | |
| ## License | |
| Derived from Qwen-Image-2.1 and released under the **Qwen Research License Agreement** (see `LICENSE`), i.e. for | |
| research / non-commercial use under its terms. The text encoder weights come from Qwen3-VL-8B-Instruct (Apache-2.0). | |