evankuo commited on
Commit
889510b
·
verified ·
1 Parent(s): ed6891e

Add files using upload-large-folder tool

Browse files
Files changed (3) hide show
  1. README.md +11 -4
  2. dit.mnn +2 -2
  3. dit.mnn.weight +1 -1
README.md CHANGED
@@ -18,14 +18,16 @@ tags:
18
  # Qwen-Image-2.1 · MNN (int4) for Android
19
 
20
  [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) converted to [MNN](https://github.com/alibaba/MNN)
21
- for on-device **text-to-image and image editing** on Android at about 512×512 pixels (any aspect ratio with sides that
22
- are multiples of 32, e.g. 512×512, 576×448, 608×416, 672×384), with the OpenCL GPU running the DiT.
 
23
 
24
  Runtime, Android library and demo app: **[github.com/scsonic/libQwenImage21](https://github.com/scsonic/libQwenImage21)**
25
 
26
  | Tested on | Snapdragon 8 Gen 2 (Adreno 740), 16 GB RAM, Android 13 |
27
  |---|---|
28
- | 512×512, 20 steps | ~490 s total (DiT ~19.7 s/step on OpenCL fp16) |
 
29
 
30
  ## Files
31
 
@@ -53,7 +55,12 @@ hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21
53
  - **VAE**: the residual stream is divided by 256 (exact, power of two) and RMSNorm pre-divides by max|x| so the decoder
54
  fits fp16 (it peaks at ~3.5e5 otherwise). The decoded image is unchanged.
55
  - The pipeline caches the text K/V once per prompt (Qwen-Image-2.1's block-causal attention), so each denoising step
56
- only runs the 1024 image tokens.
 
 
 
 
 
57
 
58
  Conversion scripts: `export/` in the GitHub repo.
59
 
 
18
  # Qwen-Image-2.1 · MNN (int4) for Android
19
 
20
  [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) converted to [MNN](https://github.com/alibaba/MNN)
21
+ for on-device **text-to-image and image editing** on Android, with the OpenCL GPU running the DiT. Any size with sides
22
+ a multiple of 32 works, from 256×256 up; the app offers 7 aspect ratios at three pixel budgets (~512², ~384², ~320²),
23
+ e.g. 512×512, 576×448, 672×384, 480×320, 384×288.
24
 
25
  Runtime, Android library and demo app: **[github.com/scsonic/libQwenImage21](https://github.com/scsonic/libQwenImage21)**
26
 
27
  | Tested on | Snapdragon 8 Gen 2 (Adreno 740), 16 GB RAM, Android 13 |
28
  |---|---|
29
+ | Text to image, 448×576, 20 steps | 451 s total (DiT 19.1 s/step on OpenCL fp16) |
30
+ | Image edit, 352×448, 20 steps | 348 s total (DiT 12.6 s/step) |
31
 
32
  ## Files
33
 
 
55
  - **VAE**: the residual stream is divided by 256 (exact, power of two) and RMSNorm pre-divides by max|x| so the decoder
56
  fits fp16 (it peaks at ~3.5e5 otherwise). The decoded image is unchanged.
57
  - The pipeline caches the text K/V once per prompt (Qwen-Image-2.1's block-causal attention), so each denoising step
58
+ only runs the image tokens. The cache is one tensor per layer (`past_kv_0`…`past_kv_31`): a single
59
+ `[32, 2, P, 32, 128]` tensor is exactly 1 MiB per prefix token, and an image-edit prefix (P > 1000) would exceed
60
+ OpenCL's 1 GiB maximum buffer size on an Adreno 740.
61
+
62
+ > **2026-09-23:** `dit.mnn` / `dit.mnn.weight` were re-exported for that per-layer K/V cache. Older copies do not load
63
+ > with the current runtime — re-download both files.
64
 
65
  Conversion scripts: `export/` in the GitHub repo.
66
 
dit.mnn CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:06afa72d2e180a369de30dc534eafbfa6796bc505bdade22d9c89ff0fe68d721
3
- size 756880
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:069f8d9fe3621a26821902eed7a07d98af5387df298f2ffc664817cc0d2ad90f
3
+ size 700184
dit.mnn.weight CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:9e74678793b4b8d30bd82e5bc01ce59dc498748ec273a7f95ad418d1c8114c6a
3
  size 4472625758
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0662e992a31c59b4ecd0a044e612877c058e5390d64c9d71b19313278f7f746a
3
  size 4472625758