Add files using upload-large-folder tool
Browse files- README.md +11 -4
- dit.mnn +2 -2
- dit.mnn.weight +1 -1
README.md
CHANGED
|
@@ -18,14 +18,16 @@ tags:
|
|
| 18 |
# Qwen-Image-2.1 · MNN (int4) for Android
|
| 19 |
|
| 20 |
[Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) converted to [MNN](https://github.com/alibaba/MNN)
|
| 21 |
-
for on-device **text-to-image and image editing** on Android
|
| 22 |
-
|
|
|
|
| 23 |
|
| 24 |
Runtime, Android library and demo app: **[github.com/scsonic/libQwenImage21](https://github.com/scsonic/libQwenImage21)**
|
| 25 |
|
| 26 |
| Tested on | Snapdragon 8 Gen 2 (Adreno 740), 16 GB RAM, Android 13 |
|
| 27 |
|---|---|
|
| 28 |
-
|
|
|
|
|
| 29 |
|
| 30 |
## Files
|
| 31 |
|
|
@@ -53,7 +55,12 @@ hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21
|
|
| 53 |
- **VAE**: the residual stream is divided by 256 (exact, power of two) and RMSNorm pre-divides by max|x| so the decoder
|
| 54 |
fits fp16 (it peaks at ~3.5e5 otherwise). The decoded image is unchanged.
|
| 55 |
- The pipeline caches the text K/V once per prompt (Qwen-Image-2.1's block-causal attention), so each denoising step
|
| 56 |
-
only runs the
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 57 |
|
| 58 |
Conversion scripts: `export/` in the GitHub repo.
|
| 59 |
|
|
|
|
| 18 |
# Qwen-Image-2.1 · MNN (int4) for Android
|
| 19 |
|
| 20 |
[Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) converted to [MNN](https://github.com/alibaba/MNN)
|
| 21 |
+
for on-device **text-to-image and image editing** on Android, with the OpenCL GPU running the DiT. Any size with sides
|
| 22 |
+
a multiple of 32 works, from 256×256 up; the app offers 7 aspect ratios at three pixel budgets (~512², ~384², ~320²),
|
| 23 |
+
e.g. 512×512, 576×448, 672×384, 480×320, 384×288.
|
| 24 |
|
| 25 |
Runtime, Android library and demo app: **[github.com/scsonic/libQwenImage21](https://github.com/scsonic/libQwenImage21)**
|
| 26 |
|
| 27 |
| Tested on | Snapdragon 8 Gen 2 (Adreno 740), 16 GB RAM, Android 13 |
|
| 28 |
|---|---|
|
| 29 |
+
| Text to image, 448×576, 20 steps | 451 s total (DiT 19.1 s/step on OpenCL fp16) |
|
| 30 |
+
| Image edit, 352×448, 20 steps | 348 s total (DiT 12.6 s/step) |
|
| 31 |
|
| 32 |
## Files
|
| 33 |
|
|
|
|
| 55 |
- **VAE**: the residual stream is divided by 256 (exact, power of two) and RMSNorm pre-divides by max|x| so the decoder
|
| 56 |
fits fp16 (it peaks at ~3.5e5 otherwise). The decoded image is unchanged.
|
| 57 |
- The pipeline caches the text K/V once per prompt (Qwen-Image-2.1's block-causal attention), so each denoising step
|
| 58 |
+
only runs the image tokens. The cache is one tensor per layer (`past_kv_0`…`past_kv_31`): a single
|
| 59 |
+
`[32, 2, P, 32, 128]` tensor is exactly 1 MiB per prefix token, and an image-edit prefix (P > 1000) would exceed
|
| 60 |
+
OpenCL's 1 GiB maximum buffer size on an Adreno 740.
|
| 61 |
+
|
| 62 |
+
> **2026-09-23:** `dit.mnn` / `dit.mnn.weight` were re-exported for that per-layer K/V cache. Older copies do not load
|
| 63 |
+
> with the current runtime — re-download both files.
|
| 64 |
|
| 65 |
Conversion scripts: `export/` in the GitHub repo.
|
| 66 |
|
dit.mnn
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:069f8d9fe3621a26821902eed7a07d98af5387df298f2ffc664817cc0d2ad90f
|
| 3 |
+
size 700184
|
dit.mnn.weight
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
size 4472625758
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:0662e992a31c59b4ecd0a044e612877c058e5390d64c9d71b19313278f7f746a
|
| 3 |
size 4472625758
|