Upload folder using huggingface_hub
Browse files- .gitattributes +8 -0
- README.md +59 -3
- samples/sample_coffee_shop.png +3 -0
- samples/sample_fisherman_448x576.png +3 -0
- samples/turbo_edit_office.png +3 -0
- samples/turbo_edit_seifuku.png +3 -0
- samples/turbo_edit_yukata.png +3 -0
- samples/turbo_t2i_harajuku.png +3 -0
- samples/turbo_t2i_kimono.png +3 -0
- samples/turbo_t2i_studio.png +3 -0
.gitattributes
CHANGED
|
@@ -45,3 +45,11 @@ vae_encoder.mnn filter=lfs diff=lfs merge=lfs -text
|
|
| 45 |
text_encoder/visual.mnn.weight filter=lfs diff=lfs merge=lfs -text
|
| 46 |
dit_turbo.mnn filter=lfs diff=lfs merge=lfs -text
|
| 47 |
dit_turbo.mnn.weight filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 45 |
text_encoder/visual.mnn.weight filter=lfs diff=lfs merge=lfs -text
|
| 46 |
dit_turbo.mnn filter=lfs diff=lfs merge=lfs -text
|
| 47 |
dit_turbo.mnn.weight filter=lfs diff=lfs merge=lfs -text
|
| 48 |
+
samples/sample_coffee_shop.png filter=lfs diff=lfs merge=lfs -text
|
| 49 |
+
samples/sample_fisherman_448x576.png filter=lfs diff=lfs merge=lfs -text
|
| 50 |
+
samples/turbo_edit_office.png filter=lfs diff=lfs merge=lfs -text
|
| 51 |
+
samples/turbo_edit_seifuku.png filter=lfs diff=lfs merge=lfs -text
|
| 52 |
+
samples/turbo_edit_yukata.png filter=lfs diff=lfs merge=lfs -text
|
| 53 |
+
samples/turbo_t2i_harajuku.png filter=lfs diff=lfs merge=lfs -text
|
| 54 |
+
samples/turbo_t2i_kimono.png filter=lfs diff=lfs merge=lfs -text
|
| 55 |
+
samples/turbo_t2i_studio.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -20,31 +20,84 @@ tags:
|
|
| 20 |
[Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) converted to [MNN](https://github.com/alibaba/MNN)
|
| 21 |
for on-device **text-to-image and image editing** on Android, with the OpenCL GPU running the DiT. Any size with sides
|
| 22 |
a multiple of 32 works, from 256Γ256 up; the app offers 7 aspect ratios at three pixel budgets (~512Β², ~384Β², ~320Β²),
|
| 23 |
-
e.g. 512Γ512, 576Γ448, 672Γ384, 480Γ320, 384Γ288.
|
|
|
|
| 24 |
|
| 25 |
Runtime, Android library and demo app: **[github.com/scsonic/libQwenImage21](https://github.com/scsonic/libQwenImage21)**
|
| 26 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 27 |
| Tested on | Snapdragon 8 Gen 2 (Adreno 740), 16 GB RAM, Android 13 |
|
| 28 |
|---|---|
|
| 29 |
| Text to image, 448Γ576, 20 steps | 451 s total (DiT 19.1 s/step on OpenCL fp16) |
|
| 30 |
| Image edit, 352Γ448, 20 steps | 348 s total (DiT 12.6 s/step) |
|
|
|
|
|
|
|
| 31 |
|
| 32 |
## Files
|
| 33 |
|
| 34 |
| Path | What | Size |
|
| 35 |
|---|---|---|
|
| 36 |
| `dit.mnn` + `.weight` | 7B single-stream DiT (32 blocks + norm_out/proj_out). Block linears int4 (block 32), small layers int8 | 4.5 GB |
|
|
|
|
| 37 |
| `txt_in.mnn`, `img_in.mnn` | text (int8) / latent (fp16) input projections | 36 MB |
|
| 38 |
| `vae_decoder.mnn` | VAE decoder, 64-ch latent β RGBA, fp16 weights, dynamic size | 0.5 GB |
|
| 39 |
| `vae_encoder.mnn` | VAE encoder for image editing, RGBA β normalized 64-ch latent, fp16 | 0.16 GB |
|
| 40 |
| `text_encoder/` | Qwen3-VL-8B-Instruct, MNN int4 (from [taobao-mnn/Qwen3-VL-8B-Instruct-MNN](https://huggingface.co/taobao-mnn/Qwen3-VL-8B-Instruct-MNN)). `te_config.json` runs it text-only; `te_vl_config.json` adds the vision tower (`visual.mnn`) for image editing. Both return the last decoder layer before the final norm | 5.4 GB |
|
| 41 |
|
| 42 |
-
Download:
|
| 43 |
|
| 44 |
```bash
|
| 45 |
-
hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21
|
|
|
|
| 46 |
```
|
| 47 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 48 |
## How it was made
|
| 49 |
|
| 50 |
- **DiT**: taken from the GGUF **Q4_K** build ([leejet/Qwen-Image-2.1-GGUF](https://huggingface.co/leejet/Qwen-Image-2.1-GGUF)).
|
|
@@ -61,6 +114,7 @@ hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21
|
|
| 61 |
|
| 62 |
> **2026-09-23:** `dit.mnn` / `dit.mnn.weight` were re-exported for that per-layer K/V cache. Older copies do not load
|
| 63 |
> with the current runtime β re-download both files.
|
|
|
|
| 64 |
|
| 65 |
Conversion scripts: `export/` in the GitHub repo.
|
| 66 |
|
|
@@ -68,3 +122,5 @@ Conversion scripts: `export/` in the GitHub repo.
|
|
| 68 |
|
| 69 |
Derived from Qwen-Image-2.1 and released under the **Qwen Research License Agreement** (see `LICENSE`), i.e. for
|
| 70 |
research / non-commercial use under its terms. The text encoder weights come from Qwen3-VL-8B-Instruct (Apache-2.0).
|
|
|
|
|
|
|
|
|
| 20 |
[Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) converted to [MNN](https://github.com/alibaba/MNN)
|
| 21 |
for on-device **text-to-image and image editing** on Android, with the OpenCL GPU running the DiT. Any size with sides
|
| 22 |
a multiple of 32 works, from 256Γ256 up; the app offers 7 aspect ratios at three pixel budgets (~512Β², ~384Β², ~320Β²),
|
| 23 |
+
e.g. 512Γ512, 576Γ448, 672Γ384, 480Γ320, 384Γ288. An optional **Turbo** LoRA (`dit_turbo.mnn`) cuts a run from
|
| 24 |
+
20β40 steps to a fixed 6.
|
| 25 |
|
| 26 |
Runtime, Android library and demo app: **[github.com/scsonic/libQwenImage21](https://github.com/scsonic/libQwenImage21)**
|
| 27 |
|
| 28 |
+
<p>
|
| 29 |
+
<img src="samples/sample_coffee_shop.png" width="24%"/>
|
| 30 |
+
<img src="samples/sample_fisherman_448x576.png" width="19%"/>
|
| 31 |
+
<img src="samples/turbo_t2i_kimono.png" width="19%"/>
|
| 32 |
+
<img src="samples/turbo_edit_yukata.png" width="19%"/>
|
| 33 |
+
<img src="samples/turbo_edit_office.png" width="19%"/>
|
| 34 |
+
</p>
|
| 35 |
+
|
| 36 |
+
*Left two: base model, 20 steps. Right three: Turbo, 6 steps β a text-to-image portrait, then two edits of it (same
|
| 37 |
+
face, new outfit and background). All generated on a phone; see [Turbo (6-step)](#turbo-6-step) below for the rest
|
| 38 |
+
of the set and per-run timings.*
|
| 39 |
+
|
| 40 |
| Tested on | Snapdragon 8 Gen 2 (Adreno 740), 16 GB RAM, Android 13 |
|
| 41 |
|---|---|
|
| 42 |
| Text to image, 448Γ576, 20 steps | 451 s total (DiT 19.1 s/step on OpenCL fp16) |
|
| 43 |
| Image edit, 352Γ448, 20 steps | 348 s total (DiT 12.6 s/step) |
|
| 44 |
+
| Text to image, 512Γ512, **Turbo 6 steps** | ~216β235 s total (DiT ~22 s/step) |
|
| 45 |
+
| Image edit, 352Γ448, **Turbo 6 steps** | ~196β200 s total |
|
| 46 |
|
| 47 |
## Files
|
| 48 |
|
| 49 |
| Path | What | Size |
|
| 50 |
|---|---|---|
|
| 51 |
| `dit.mnn` + `.weight` | 7B single-stream DiT (32 blocks + norm_out/proj_out). Block linears int4 (block 32), small layers int8 | 4.5 GB |
|
| 52 |
+
| `dit_turbo.mnn` + `.weight` | Same DiT with the [Viggle-turbo](https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo) LoRA applied unmerged (a small extra fp16 branch per targeted layer); fixed 6-step schedule. Optional β see [below](#turbo-6-step) | 5.2 GB |
|
| 53 |
| `txt_in.mnn`, `img_in.mnn` | text (int8) / latent (fp16) input projections | 36 MB |
|
| 54 |
| `vae_decoder.mnn` | VAE decoder, 64-ch latent β RGBA, fp16 weights, dynamic size | 0.5 GB |
|
| 55 |
| `vae_encoder.mnn` | VAE encoder for image editing, RGBA β normalized 64-ch latent, fp16 | 0.16 GB |
|
| 56 |
| `text_encoder/` | Qwen3-VL-8B-Instruct, MNN int4 (from [taobao-mnn/Qwen3-VL-8B-Instruct-MNN](https://huggingface.co/taobao-mnn/Qwen3-VL-8B-Instruct-MNN)). `te_config.json` runs it text-only; `te_vl_config.json` adds the vision tower (`visual.mnn`) for image editing. Both return the last decoder layer before the final norm | 5.4 GB |
|
| 57 |
|
| 58 |
+
Download everything (~15.5 GB with Turbo) or skip `dit_turbo.mnn*` to save 5.2 GB:
|
| 59 |
|
| 60 |
```bash
|
| 61 |
+
hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21 # + Turbo
|
| 62 |
+
hf download evankuo/Qwen-Image-2.1-MNN --local-dir qwen_image21 --exclude "dit_turbo.mnn*" # base model only
|
| 63 |
```
|
| 64 |
|
| 65 |
+
## Turbo (6-step)
|
| 66 |
+
|
| 67 |
+
[Viggle-turbo](https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo) distills Qwen-Image-2.1 to a fixed
|
| 68 |
+
6-step schedule with no CFG. `dit_turbo.mnn` applies it **unmerged**, the way diffusers and the LoRA's own
|
| 69 |
+
ComfyUI node do: the int4 base weights are untouched (identical to `dit.mnn`'s), and the LoRA's rank-128
|
| 70 |
+
correction is added as a small extra fp16 branch per targeted layer, exported as its own MNN model file that
|
| 71 |
+
happens to carry a second copy of the base weights (MNN has no format for patching an already-compiled graph).
|
| 72 |
+
Merging the correction into the weights instead β especially into int4 β is what the LoRA's own README
|
| 73 |
+
specifically measures as lossy; unmerged is the accurate path. `txt_in.mnn`, `img_in.mnn`, both VAE models and
|
| 74 |
+
the text encoder are unaffected and shared with the base pipeline.
|
| 75 |
+
|
| 76 |
+
<p>
|
| 77 |
+
<img src="samples/turbo_t2i_studio.png" width="19%"/>
|
| 78 |
+
<img src="samples/turbo_t2i_kimono.png" width="19%"/>
|
| 79 |
+
<img src="samples/turbo_t2i_harajuku.png" width="19%"/>
|
| 80 |
+
<img src="samples/turbo_edit_seifuku.png" width="19%"/>
|
| 81 |
+
<img src="samples/turbo_edit_yukata.png" width="19%"/>
|
| 82 |
+
</p>
|
| 83 |
+
|
| 84 |
+
*Text-to-image (studio portrait, kimono, Harajuku street fashion) and two edits of the studio portrait β same
|
| 85 |
+
face, new outfit and setting. All on a Snapdragon 8 Gen 2, OpenCL, 6 steps.*
|
| 86 |
+
|
| 87 |
+
| | text encoder (+ prefix) | DiT (6 steps) | VAE | total |
|
| 88 |
+
|---|---|---|---|---|
|
| 89 |
+
| Text to image, 448Γ576β512Γ512 | ~15 s | 6 Γ ~22 s β 134 s | ~18 s | **216β235 s** |
|
| 90 |
+
| Image edit β 352Γ448 | ~20β80 s (with vision; Pβ680) | 6 Γ ~18 s β 108 s | ~13 s | **196β200 s** |
|
| 91 |
+
|
| 92 |
+
A Turbo DiT step (~22 s at ~512Β²) is slower than a base-model step at the same size (~19 s) β the extra fp16
|
| 93 |
+
branch costs roughly the 10β25% diffusers/ComfyUI themselves measure β but 6 steps instead of 20 still roughly
|
| 94 |
+
halves the total time for both modes. Base model at 6 steps *without* the LoRA is visibly worse (soft, muddy) β
|
| 95 |
+
the schedule alone isn't what's doing the work.
|
| 96 |
+
|
| 97 |
+
Load it like the base model, just with `dit_turbo.mnn` instead of `dit.mnn`; the demo app has a **Turbo LoRA**
|
| 98 |
+
checkbox that fixes the step count to 6. See the [runtime repo](https://github.com/scsonic/libQwenImage21/blob/main/docs/TURBO.md)
|
| 99 |
+
for the CLI/library API.
|
| 100 |
+
|
| 101 |
## How it was made
|
| 102 |
|
| 103 |
- **DiT**: taken from the GGUF **Q4_K** build ([leejet/Qwen-Image-2.1-GGUF](https://huggingface.co/leejet/Qwen-Image-2.1-GGUF)).
|
|
|
|
| 114 |
|
| 115 |
> **2026-09-23:** `dit.mnn` / `dit.mnn.weight` were re-exported for that per-layer K/V cache. Older copies do not load
|
| 116 |
> with the current runtime β re-download both files.
|
| 117 |
+
> **2026-09-25:** added `dit_turbo.mnn` / `.weight` (optional, see [Turbo](#turbo-6-step) above).
|
| 118 |
|
| 119 |
Conversion scripts: `export/` in the GitHub repo.
|
| 120 |
|
|
|
|
| 122 |
|
| 123 |
Derived from Qwen-Image-2.1 and released under the **Qwen Research License Agreement** (see `LICENSE`), i.e. for
|
| 124 |
research / non-commercial use under its terms. The text encoder weights come from Qwen3-VL-8B-Instruct (Apache-2.0).
|
| 125 |
+
The Turbo LoRA is a derivative of the same base model, released by Viggle under the same Qwen Research License
|
| 126 |
+
Agreement.
|
samples/sample_coffee_shop.png
ADDED
|
Git LFS Details
|
samples/sample_fisherman_448x576.png
ADDED
|
Git LFS Details
|
samples/turbo_edit_office.png
ADDED
|
Git LFS Details
|
samples/turbo_edit_seifuku.png
ADDED
|
Git LFS Details
|
samples/turbo_edit_yukata.png
ADDED
|
Git LFS Details
|
samples/turbo_t2i_harajuku.png
ADDED
|
Git LFS Details
|
samples/turbo_t2i_kimono.png
ADDED
|
Git LFS Details
|
samples/turbo_t2i_studio.png
ADDED
|
Git LFS Details
|