evankuo commited on
Commit
2f712ce
Β·
verified Β·
1 Parent(s): 75469e9

Upload folder using huggingface_hub

Browse files
.gitattributes CHANGED
@@ -45,3 +45,11 @@ vae_encoder.mnn filter=lfs diff=lfs merge=lfs -text
45
  text_encoder/visual.mnn.weight filter=lfs diff=lfs merge=lfs -text
46
  dit_turbo.mnn filter=lfs diff=lfs merge=lfs -text
47
  dit_turbo.mnn.weight filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
 
 
45
  text_encoder/visual.mnn.weight filter=lfs diff=lfs merge=lfs -text
46
  dit_turbo.mnn filter=lfs diff=lfs merge=lfs -text
47
  dit_turbo.mnn.weight filter=lfs diff=lfs merge=lfs -text
48
+ samples/sample_coffee_shop.png filter=lfs diff=lfs merge=lfs -text
49
+ samples/sample_fisherman_448x576.png filter=lfs diff=lfs merge=lfs -text
50
+ samples/turbo_edit_office.png filter=lfs diff=lfs merge=lfs -text
51
+ samples/turbo_edit_seifuku.png filter=lfs diff=lfs merge=lfs -text
52
+ samples/turbo_edit_yukata.png filter=lfs diff=lfs merge=lfs -text
53
+ samples/turbo_t2i_harajuku.png filter=lfs diff=lfs merge=lfs -text
54
+ samples/turbo_t2i_kimono.png filter=lfs diff=lfs merge=lfs -text
55
+ samples/turbo_t2i_studio.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -20,31 +20,84 @@ tags:
20
  [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) converted to [MNN](https://github.com/alibaba/MNN)
21
  for on-device **text-to-image and image editing** on Android, with the OpenCL GPU running the DiT. Any size with sides
22
  a multiple of 32 works, from 256Γ—256 up; the app offers 7 aspect ratios at three pixel budgets (~512Β², ~384Β², ~320Β²),
23
- e.g. 512Γ—512, 576Γ—448, 672Γ—384, 480Γ—320, 384Γ—288.
 
24
 
25
  Runtime, Android library and demo app: **[github.com/scsonic/libQwenImage21](https://github.com/scsonic/libQwenImage21)**
26
 
 
 
 
 
 
 
 
 
 
 
 
 
27
  | Tested on | Snapdragon 8 Gen 2 (Adreno 740), 16 GB RAM, Android 13 |
28
  |---|---|
29
  | Text to image, 448Γ—576, 20 steps | 451 s total (DiT 19.1 s/step on OpenCL fp16) |
30
  | Image edit, 352Γ—448, 20 steps | 348 s total (DiT 12.6 s/step) |
 
 
31
 
32
  ## Files
33
 
34
  | Path | What | Size |
35
  |---|---|---|
36
  | `dit.mnn` + `.weight` | 7B single-stream DiT (32 blocks + norm_out/proj_out). Block linears int4 (block 32), small layers int8 | 4.5 GB |
 
37
  | `txt_in.mnn`, `img_in.mnn` | text (int8) / latent (fp16) input projections | 36 MB |
38
  | `vae_decoder.mnn` | VAE decoder, 64-ch latent β†’ RGBA, fp16 weights, dynamic size | 0.5 GB |
39
  | `vae_encoder.mnn` | VAE encoder for image editing, RGBA β†’ normalized 64-ch latent, fp16 | 0.16 GB |
40
  | `text_encoder/` | Qwen3-VL-8B-Instruct, MNN int4 (from [taobao-mnn/Qwen3-VL-8B-Instruct-MNN](https://huggingface.co/taobao-mnn/Qwen3-VL-8B-Instruct-MNN)). `te_config.json` runs it text-only; `te_vl_config.json` adds the vision tower (`visual.mnn`) for image editing. Both return the last decoder layer before the final norm | 5.4 GB |
41
 
42
- Download:
43
 
44
  ```bash
45
- hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21
 
46
  ```
47
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
48
  ## How it was made
49
 
50
  - **DiT**: taken from the GGUF **Q4_K** build ([leejet/Qwen-Image-2.1-GGUF](https://huggingface.co/leejet/Qwen-Image-2.1-GGUF)).
@@ -61,6 +114,7 @@ hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21
61
 
62
  > **2026-09-23:** `dit.mnn` / `dit.mnn.weight` were re-exported for that per-layer K/V cache. Older copies do not load
63
  > with the current runtime β€” re-download both files.
 
64
 
65
  Conversion scripts: `export/` in the GitHub repo.
66
 
@@ -68,3 +122,5 @@ Conversion scripts: `export/` in the GitHub repo.
68
 
69
  Derived from Qwen-Image-2.1 and released under the **Qwen Research License Agreement** (see `LICENSE`), i.e. for
70
  research / non-commercial use under its terms. The text encoder weights come from Qwen3-VL-8B-Instruct (Apache-2.0).
 
 
 
20
  [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) converted to [MNN](https://github.com/alibaba/MNN)
21
  for on-device **text-to-image and image editing** on Android, with the OpenCL GPU running the DiT. Any size with sides
22
  a multiple of 32 works, from 256Γ—256 up; the app offers 7 aspect ratios at three pixel budgets (~512Β², ~384Β², ~320Β²),
23
+ e.g. 512Γ—512, 576Γ—448, 672Γ—384, 480Γ—320, 384Γ—288. An optional **Turbo** LoRA (`dit_turbo.mnn`) cuts a run from
24
+ 20–40 steps to a fixed 6.
25
 
26
  Runtime, Android library and demo app: **[github.com/scsonic/libQwenImage21](https://github.com/scsonic/libQwenImage21)**
27
 
28
+ <p>
29
+ <img src="samples/sample_coffee_shop.png" width="24%"/>
30
+ <img src="samples/sample_fisherman_448x576.png" width="19%"/>
31
+ <img src="samples/turbo_t2i_kimono.png" width="19%"/>
32
+ <img src="samples/turbo_edit_yukata.png" width="19%"/>
33
+ <img src="samples/turbo_edit_office.png" width="19%"/>
34
+ </p>
35
+
36
+ *Left two: base model, 20 steps. Right three: Turbo, 6 steps β€” a text-to-image portrait, then two edits of it (same
37
+ face, new outfit and background). All generated on a phone; see [Turbo (6-step)](#turbo-6-step) below for the rest
38
+ of the set and per-run timings.*
39
+
40
  | Tested on | Snapdragon 8 Gen 2 (Adreno 740), 16 GB RAM, Android 13 |
41
  |---|---|
42
  | Text to image, 448Γ—576, 20 steps | 451 s total (DiT 19.1 s/step on OpenCL fp16) |
43
  | Image edit, 352Γ—448, 20 steps | 348 s total (DiT 12.6 s/step) |
44
+ | Text to image, 512Γ—512, **Turbo 6 steps** | ~216–235 s total (DiT ~22 s/step) |
45
+ | Image edit, 352Γ—448, **Turbo 6 steps** | ~196–200 s total |
46
 
47
  ## Files
48
 
49
  | Path | What | Size |
50
  |---|---|---|
51
  | `dit.mnn` + `.weight` | 7B single-stream DiT (32 blocks + norm_out/proj_out). Block linears int4 (block 32), small layers int8 | 4.5 GB |
52
+ | `dit_turbo.mnn` + `.weight` | Same DiT with the [Viggle-turbo](https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo) LoRA applied unmerged (a small extra fp16 branch per targeted layer); fixed 6-step schedule. Optional β€” see [below](#turbo-6-step) | 5.2 GB |
53
  | `txt_in.mnn`, `img_in.mnn` | text (int8) / latent (fp16) input projections | 36 MB |
54
  | `vae_decoder.mnn` | VAE decoder, 64-ch latent β†’ RGBA, fp16 weights, dynamic size | 0.5 GB |
55
  | `vae_encoder.mnn` | VAE encoder for image editing, RGBA β†’ normalized 64-ch latent, fp16 | 0.16 GB |
56
  | `text_encoder/` | Qwen3-VL-8B-Instruct, MNN int4 (from [taobao-mnn/Qwen3-VL-8B-Instruct-MNN](https://huggingface.co/taobao-mnn/Qwen3-VL-8B-Instruct-MNN)). `te_config.json` runs it text-only; `te_vl_config.json` adds the vision tower (`visual.mnn`) for image editing. Both return the last decoder layer before the final norm | 5.4 GB |
57
 
58
+ Download everything (~15.5 GB with Turbo) or skip `dit_turbo.mnn*` to save 5.2 GB:
59
 
60
  ```bash
61
+ hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21 # + Turbo
62
+ hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21 --exclude "dit_turbo.mnn*" # base model only
63
  ```
64
 
65
+ ## Turbo (6-step)
66
+
67
+ [Viggle-turbo](https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo) distills Qwen-Image-2.1 to a fixed
68
+ 6-step schedule with no CFG. `dit_turbo.mnn` applies it **unmerged**, the way diffusers and the LoRA's own
69
+ ComfyUI node do: the int4 base weights are untouched (identical to `dit.mnn`'s), and the LoRA's rank-128
70
+ correction is added as a small extra fp16 branch per targeted layer, exported as its own MNN model file that
71
+ happens to carry a second copy of the base weights (MNN has no format for patching an already-compiled graph).
72
+ Merging the correction into the weights instead β€” especially into int4 β€” is what the LoRA's own README
73
+ specifically measures as lossy; unmerged is the accurate path. `txt_in.mnn`, `img_in.mnn`, both VAE models and
74
+ the text encoder are unaffected and shared with the base pipeline.
75
+
76
+ <p>
77
+ <img src="samples/turbo_t2i_studio.png" width="19%"/>
78
+ <img src="samples/turbo_t2i_kimono.png" width="19%"/>
79
+ <img src="samples/turbo_t2i_harajuku.png" width="19%"/>
80
+ <img src="samples/turbo_edit_seifuku.png" width="19%"/>
81
+ <img src="samples/turbo_edit_yukata.png" width="19%"/>
82
+ </p>
83
+
84
+ *Text-to-image (studio portrait, kimono, Harajuku street fashion) and two edits of the studio portrait β€” same
85
+ face, new outfit and setting. All on a Snapdragon 8 Gen 2, OpenCL, 6 steps.*
86
+
87
+ | | text encoder (+ prefix) | DiT (6 steps) | VAE | total |
88
+ |---|---|---|---|---|
89
+ | Text to image, 448Γ—576–512Γ—512 | ~15 s | 6 Γ— ~22 s β‰ˆ 134 s | ~18 s | **216–235 s** |
90
+ | Image edit β†’ 352Γ—448 | ~20–80 s (with vision; Pβ‰ˆ680) | 6 Γ— ~18 s β‰ˆ 108 s | ~13 s | **196–200 s** |
91
+
92
+ A Turbo DiT step (~22 s at ~512Β²) is slower than a base-model step at the same size (~19 s) β€” the extra fp16
93
+ branch costs roughly the 10–25% diffusers/ComfyUI themselves measure β€” but 6 steps instead of 20 still roughly
94
+ halves the total time for both modes. Base model at 6 steps *without* the LoRA is visibly worse (soft, muddy) β€”
95
+ the schedule alone isn't what's doing the work.
96
+
97
+ Load it like the base model, just with `dit_turbo.mnn` instead of `dit.mnn`; the demo app has a **Turbo LoRA**
98
+ checkbox that fixes the step count to 6. See the [runtime repo](https://github.com/scsonic/libQwenImage21/blob/main/docs/TURBO.md)
99
+ for the CLI/library API.
100
+
101
  ## How it was made
102
 
103
  - **DiT**: taken from the GGUF **Q4_K** build ([leejet/Qwen-Image-2.1-GGUF](https://huggingface.co/leejet/Qwen-Image-2.1-GGUF)).
 
114
 
115
  > **2026-09-23:** `dit.mnn` / `dit.mnn.weight` were re-exported for that per-layer K/V cache. Older copies do not load
116
  > with the current runtime β€” re-download both files.
117
+ > **2026-09-25:** added `dit_turbo.mnn` / `.weight` (optional, see [Turbo](#turbo-6-step) above).
118
 
119
  Conversion scripts: `export/` in the GitHub repo.
120
 
 
122
 
123
  Derived from Qwen-Image-2.1 and released under the **Qwen Research License Agreement** (see `LICENSE`), i.e. for
124
  research / non-commercial use under its terms. The text encoder weights come from Qwen3-VL-8B-Instruct (Apache-2.0).
125
+ The Turbo LoRA is a derivative of the same base model, released by Viggle under the same Qwen Research License
126
+ Agreement.
samples/sample_coffee_shop.png ADDED

Git LFS Details

  • SHA256: e76d81fa7d97f2a86c481ef639b70de05eafe6ed6ba285a384035a94dfd2949a
  • Pointer size: 131 Bytes
  • Size of remote file: 438 kB
samples/sample_fisherman_448x576.png ADDED

Git LFS Details

  • SHA256: 72eb1e0e436533436f4d9f7f038e5a142ad6eadb82d7a3510cf98dc9a4a02c74
  • Pointer size: 131 Bytes
  • Size of remote file: 478 kB
samples/turbo_edit_office.png ADDED

Git LFS Details

  • SHA256: 1d6d1bea85f68395e2eefc3b1f468d692b1b50fec820e872137010a629fb240b
  • Pointer size: 131 Bytes
  • Size of remote file: 235 kB
samples/turbo_edit_seifuku.png ADDED

Git LFS Details

  • SHA256: 982daa7a966f9b2dc69c6cdea3f2433331506a739c14bd8ae00f2c647a01e0f5
  • Pointer size: 131 Bytes
  • Size of remote file: 296 kB
samples/turbo_edit_yukata.png ADDED

Git LFS Details

  • SHA256: c2428151f258b7b6a975cf42e693e1d205d00833d117eec3c9c9855ebac7fc77
  • Pointer size: 131 Bytes
  • Size of remote file: 321 kB
samples/turbo_t2i_harajuku.png ADDED

Git LFS Details

  • SHA256: 047733deab19b7701feba43f00be43060873ca942be1a57b9f249770f345fc91
  • Pointer size: 131 Bytes
  • Size of remote file: 558 kB
samples/turbo_t2i_kimono.png ADDED

Git LFS Details

  • SHA256: b96bf7de91bed80c670e896b5726b06ac51240deb6859fd19d559de04a9516f9
  • Pointer size: 131 Bytes
  • Size of remote file: 541 kB
samples/turbo_t2i_studio.png ADDED

Git LFS Details

  • SHA256: 3cd1802a5d6893ca95fe83c7888a36d2f161eea2161a54461f318e400df1417a
  • Pointer size: 131 Bytes
  • Size of remote file: 415 kB