Spaces:
Running on Zero
Running on Zero
v0.3: 6 steps by default, new 9-step setting (7 turbo + 2 base-model steps)
Browse filesThis view is limited to 50 files because it contains too many changes. See raw diff
- .gitattributes +25 -0
- NOTICE +1 -1
- README.md +37 -240
- app.py +56 -32
- compare/cases.json +0 -0
- compare/img/case00/turbo.webp +2 -2
- compare/img/case00/turbo9.webp +3 -0
- compare/img/{case01/turbo8.webp → case00/turbo9_t.webp} +2 -2
- compare/img/case00/turbo_t.webp +0 -0
- compare/img/case01/turbo.webp +2 -2
- compare/img/case01/turbo8_t.webp +0 -0
- compare/img/{case17/turbo8.webp → case01/turbo9.webp} +2 -2
- compare/img/case01/turbo9_t.webp +0 -0
- compare/img/case01/turbo_t.webp +0 -0
- compare/img/case02/turbo.webp +2 -2
- compare/img/{case16/turbo8.webp → case02/turbo9.webp} +2 -2
- compare/img/case02/turbo9_t.webp +0 -0
- compare/img/case02/turbo_t.webp +2 -2
- compare/img/case03/turbo.webp +2 -2
- compare/img/{case18/turbo8.webp → case03/turbo9.webp} +2 -2
- compare/img/case03/turbo9_t.webp +3 -0
- compare/img/case03/turbo_t.webp +2 -2
- compare/img/case04/turbo.webp +2 -2
- compare/img/case04/turbo9.webp +3 -0
- compare/img/case04/turbo9_t.webp +0 -0
- compare/img/case04/turbo_t.webp +0 -0
- compare/img/case05/turbo.webp +2 -2
- compare/img/case05/turbo9.webp +3 -0
- compare/img/case05/turbo9_t.webp +0 -0
- compare/img/case05/turbo_t.webp +0 -0
- compare/img/case06/turbo.webp +2 -2
- compare/img/case06/turbo9.webp +3 -0
- compare/img/case06/turbo9_t.webp +0 -0
- compare/img/case06/turbo_t.webp +2 -2
- compare/img/case07/turbo.webp +2 -2
- compare/img/case07/turbo9.webp +3 -0
- compare/img/case07/turbo9_t.webp +0 -0
- compare/img/case07/turbo_t.webp +0 -0
- compare/img/case08/turbo.webp +2 -2
- compare/img/case08/turbo9.webp +3 -0
- compare/img/case08/turbo9_t.webp +0 -0
- compare/img/case08/turbo_t.webp +0 -0
- compare/img/case09/turbo.webp +2 -2
- compare/img/case09/turbo9.webp +3 -0
- compare/img/case09/turbo9_t.webp +0 -0
- compare/img/case09/turbo_t.webp +0 -0
- compare/img/case10/turbo.webp +2 -2
- compare/img/case10/turbo9.webp +3 -0
- compare/img/case10/turbo9_t.webp +0 -0
- compare/img/case10/turbo_t.webp +0 -0
.gitattributes
CHANGED
|
@@ -197,3 +197,28 @@ examples/qwen_rgba_bride.png filter=lfs diff=lfs merge=lfs -text
|
|
| 197 |
examples/qwen_storyboard.png filter=lfs diff=lfs merge=lfs -text
|
| 198 |
examples/qwen_storyboard_in1.webp filter=lfs diff=lfs merge=lfs -text
|
| 199 |
examples/t2i_style_file.png filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 197 |
examples/qwen_storyboard.png filter=lfs diff=lfs merge=lfs -text
|
| 198 |
examples/qwen_storyboard_in1.webp filter=lfs diff=lfs merge=lfs -text
|
| 199 |
examples/t2i_style_file.png filter=lfs diff=lfs merge=lfs -text
|
| 200 |
+
compare/img/case00/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 201 |
+
compare/img/case00/turbo9_t.webp filter=lfs diff=lfs merge=lfs -text
|
| 202 |
+
compare/img/case00/turbo_t.webp filter=lfs diff=lfs merge=lfs -text
|
| 203 |
+
compare/img/case01/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 204 |
+
compare/img/case02/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 205 |
+
compare/img/case03/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 206 |
+
compare/img/case03/turbo9_t.webp filter=lfs diff=lfs merge=lfs -text
|
| 207 |
+
compare/img/case04/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 208 |
+
compare/img/case05/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 209 |
+
compare/img/case06/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 210 |
+
compare/img/case07/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 211 |
+
compare/img/case08/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 212 |
+
compare/img/case09/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 213 |
+
compare/img/case10/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 214 |
+
compare/img/case11/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 215 |
+
compare/img/case12/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 216 |
+
compare/img/case13/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 217 |
+
compare/img/case14/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 218 |
+
compare/img/case15/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 219 |
+
compare/img/case16/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 220 |
+
compare/img/case17/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 221 |
+
compare/img/case18/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 222 |
+
compare/img/case19/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 223 |
+
compare/img/case20/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
| 224 |
+
compare/img/case21/turbo9.webp filter=lfs diff=lfs merge=lfs -text
|
NOTICE
CHANGED
|
@@ -2,7 +2,7 @@ Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 H
|
|
| 2 |
|
| 3 |
Built with Qwen.
|
| 4 |
|
| 5 |
-
This Space demonstrates Viggle/Qwen-Image-2.1-viggle-turbo, a
|
| 6 |
derivative of Qwen/Qwen-Image-2.1. The base weights are downloaded at runtime from
|
| 7 |
Qwen/Qwen-Image-2.1 and are unmodified. Use is limited to non-commercial (research or
|
| 8 |
evaluation) purposes, per the Qwen RESEARCH LICENSE AGREEMENT included as LICENSE.
|
|
|
|
| 2 |
|
| 3 |
Built with Qwen.
|
| 4 |
|
| 5 |
+
This Space demonstrates Viggle/Qwen-Image-2.1-viggle-turbo, a distilled few-step LoRA
|
| 6 |
derivative of Qwen/Qwen-Image-2.1. The base weights are downloaded at runtime from
|
| 7 |
Qwen/Qwen-Image-2.1 and are unmodified. Use is limited to non-commercial (research or
|
| 8 |
evaluation) purposes, per the Qwen RESEARCH LICENSE AGREEMENT included as LICENSE.
|
README.md
CHANGED
|
@@ -19,266 +19,63 @@ tags:
|
|
| 19 |
- text-to-image
|
| 20 |
- image-editing
|
| 21 |
- distillation
|
| 22 |
-
- dmd
|
| 23 |
# Hardware: this Space needs ZeroGPU, set in Space Settings. `suggested_hardware` is deliberately
|
| 24 |
# unset, because the only ZeroGPU value the Hub metadata accepts is the legacy `zero-a10g`, and a
|
| 25 |
# 24 GB A10G cannot hold this pipeline. `app.py` requests the 96 GB card with
|
| 26 |
-
# @spaces.GPU(size="xlarge") (no VAE tiling)
|
| 27 |
# Secrets: set HF_TOKEN (read scope) while Viggle/Qwen-Image-2.1-viggle-turbo is private.
|
| 28 |
# `hf_oauth` is not needed: the app calls no user-scoped Hub API.
|
| 29 |
---
|
| 30 |
|
| 31 |
-
# Viggle Turbo v0.
|
| 32 |
|
| 33 |
-
A
|
| 34 |
-
with **no classifier-free guidance**,
|
| 35 |
-
hard to tell apart from the base model; small, dense text
|
| 36 |
-
|
| 37 |
-
tab shows it side by side with the base model on the official Qwen examples.
|
| 38 |
|
| 39 |
-
**v0.
|
| 40 |
-
|
| 41 |
-
|
|
|
|
|
|
|
|
|
|
| 42 |
|
| 43 |
-
**
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
|
|
|
|
|
|
| 47 |
|
| 48 |
-
**
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
One model does both tasks, exactly as the base does: leave **References** empty for text-to-image,
|
| 53 |
-
or add one to six images (drop them on it, or use **+ Add image**) for editing / composition / style
|
| 54 |
-
transfer. The prompt refers to them as image 1, image 2, … in the order they were added, the number shown
|
| 55 |
-
on each thumbnail.
|
| 56 |
-
|
| 57 |
-
The **Examples** table below the form is pre-rendered: every row carries the output this exact
|
| 58 |
-
app produced for it (fixed seed, 6 steps, prompt sent as written unless the row's status line says
|
| 59 |
-
otherwise — the three-reference row has prompt enhancement on, and the launch-graphic row is 8 steps at the
|
| 60 |
-
3:2 2048² size), so you can see what the model
|
| 61 |
-
does without spending GPU time — click a row to load the prompt, the references, the result, and the
|
| 62 |
-
steps / size / enhancement settings and seed that produced it (with *Randomize seed* switched off), so
|
| 63 |
-
**Generate** reproduces the row (except the three-reference row, whose stored prompt came from the earlier
|
| 64 |
-
on-GPU rewriter; Generate rewrites it again with DeepSeek). `release/render_examples.py` regenerates the rows
|
| 65 |
-
and `examples/manifest.json` whenever the weights change. The reference photos
|
| 66 |
-
`woman1/cat_window/bird.webp` and the three-reference prompt come from the
|
| 67 |
-
[black-forest-labs/flux-klein-9b-kv](https://huggingface.co/spaces/black-forest-labs/flux-klein-9b-kv) Space;
|
| 68 |
-
the prompts and input images of the `qwen_*` rows come from the
|
| 69 |
-
[Qwen/Qwen-Image-2.1](https://huggingface.co/spaces/Qwen/Qwen-Image-2.1) Space; the character style-file
|
| 70 |
-
row is the image of our promo video.
|
| 71 |
-
|
| 72 |
-
The **Comparison** tab needs no GPU: 32 of the 37 examples of the Qwen/Qwen-Image-2.1 Space, pre-rendered with
|
| 73 |
-
this model in 6 steps (and in 8 steps on the 5 dense-text examples) and with the base model in 40 steps — same
|
| 74 |
-
prompt, input images and seed 42, prompt enhancement off, one sample each — shown in an image slider with the
|
| 75 |
-
render times, plus Qwen's own API output where their Space ships one. Its images (`compare/`) are built from
|
| 76 |
-
the renders by `release/build_compare_space.py` when `release/upload_space.sh` uploads the Space. Of the five
|
| 77 |
-
examples left out, four are Chinese infographics / a subtitled storyboard whose short prompts leave the on-image
|
| 78 |
-
text to prompt enhancement, and one asks for a named third-party character.
|
| 79 |
-
|
| 80 |
-
**Built with Qwen.** Distilled from [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1).
|
| 81 |
-
|
| 82 |
-
> **Known limits.** Complicated edits (multi-reference composition, face swaps, identity-preserving edits,
|
| 83 |
-
> instructions with several constraints) can still fall short of the 40-step base model, and at 6 steps very
|
| 84 |
-
> small dense text can print less cleanly (8 steps helps). We keep updating the weights as the distillation improves.
|
| 85 |
-
|
| 86 |
-
## What the app loads
|
| 87 |
-
|
| 88 |
-
The Space runs the **v0.2.1 LoRA (rank 256)** by default, loaded at runtime on top of the base transformer with
|
| 89 |
-
`pipe.load_lora_weights(STUDENT_REPO, weight_name=LORA_FILE)` and never merged (merging into bf16 keeps only ~47 %
|
| 90 |
-
of the adapter delta). `LORA_FILE` defaults to `Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r256.safetensors`; the
|
| 91 |
-
v0.2 and v0.1 adapters are still in the model repo and can be selected with it. The v0.1 **full fine-tuned transformer**
|
| 92 |
-
is still selectable with `STUDENT=full`:
|
| 93 |
-
|
| 94 |
-
```python
|
| 95 |
-
transformer = QwenImage21Transformer2DModel.from_pretrained(STUDENT_REPO, subfolder="transformer", torch_dtype=torch.bfloat16)
|
| 96 |
-
pipe = QwenImage21Pipeline.from_pretrained(BASE_MODEL_ID, transformer=transformer, dtype=torch.bfloat16)
|
| 97 |
-
```
|
| 98 |
-
|
| 99 |
-
`transformer/` is the v0.1 full-parameter fine-tune of the base transformer (bf16, 14.2 GB); passing it as a
|
| 100 |
-
component means the base transformer is never downloaded. The status line under each result names the
|
| 101 |
-
student in use.
|
| 102 |
-
|
| 103 |
-
Environment: `STUDENT` (`lora`, default, or `full`), `LORA_FILE`, `STUDENT_REPO` (default
|
| 104 |
-
`Viggle/Qwen-Image-2.1-viggle-turbo`; `LORA_REPO` is still accepted), `BASE_MODEL_ID` (default
|
| 105 |
-
`Qwen/Qwen-Image-2.1`), `HF_TOKEN`, `OPENROUTER_API_KEY` (prompt enhancement; without it every prompt is sent
|
| 106 |
-
as written). **`HF_TOKEN` is required while the model repo is private** — add it as a Space secret with read
|
| 107 |
-
scope; it can be removed once the repo is public.
|
| 108 |
-
|
| 109 |
-
## The 6-step contract
|
| 110 |
-
|
| 111 |
-
The student is sampled on a specific schedule, and the demo reproduces it exactly:
|
| 112 |
-
|
| 113 |
-
- **6 Euler steps**, raw sigma nodes `[1, 0.9375, 0.875, 0.75, 0.5, 0.25]` passed as `sigmas=` — the 4-step
|
| 114 |
-
training nodes `linspace(1, 1/4, 4)` with the first (highest-noise) segment `1 → 0.75` cut into three.
|
| 115 |
-
The composition is decided in that segment, and a single Euler step across it ghosts and drifts the
|
| 116 |
-
layout; the low-noise nodes `0.75, 0.5, 0.25` are the ones the student was trained to land on, and moving
|
| 117 |
-
them (e.g. to `0.625, 0.375`) makes every image visibly softer. The pipeline shifts whatever it is given, so these raw nodes
|
| 118 |
-
get the same resolution-dependent shift as its defaults.
|
| 119 |
-
- **Extra steps always go to the high-noise end.** The **Steps** slider (3–8) follows `raw_nodes(steps)`
|
| 120 |
-
in `app.py`: the nodes `0.875, 0.75, 0.5, 0.25` are fixed and the segment `1 → 0.875` is split evenly
|
| 121 |
-
into `steps − 4` pieces (5: `1, 0.875, …`, the schedule v0.2 launched with; 6: `1, 0.9375, 0.875, …`;
|
| 122 |
-
7: `1, 0.958, 0.917, 0.875, …`, also validated; 8 untested). 4 is the training schedule (a little less
|
| 123 |
-
detail and diversity), 3 falls back to the pipeline's uniform `linspace(1, 1/3, 3)` and is off-contract.
|
| 124 |
-
- **Resolution-dependent time shift only.** `mu = calculate_shift((H/16)*(W/16))` with the base
|
| 125 |
-
scheduler's `base_image_seq_len=256, max_image_seq_len=8192, base_shift=0.5, max_shift=0.9`.
|
| 126 |
-
- **`shift_terminal` must be `null`.** The stock Qwen-Image-2.1 scheduler ships `shift_terminal: 0.02`,
|
| 127 |
-
which stretches the final sigma node to 0.02 instead of 0 and costs the last step. `app.py` rebuilds
|
| 128 |
-
the scheduler from the base config with the key cleared, so the contract holds no matter which repo
|
| 129 |
-
`BASE_MODEL_ID` points at:
|
| 130 |
-
`pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_config(pipe.scheduler.config, shift_terminal=None)`
|
| 131 |
-
- **No CFG.** `true_cfg_scale=1.0` and no negative prompt. (Qwen-Image-2.1 has no `guidance_scale`
|
| 132 |
-
argument at all; CFG needs both `true_cfg_scale > 1` and a negative prompt.)
|
| 133 |
-
- Condition images are encoded at 1024-area (`output_resolution=1024`), matching distillation.
|
| 134 |
-
|
| 135 |
-
Verified against `dmd_common.student_sigmas(..., terminal="none")` with the split nodes on a 1024²-bucket
|
| 136 |
-
text-to-image case and a 1024²-bucket edit: max absolute gap **0.0**, and the public
|
| 137 |
-
`load_lora_weights` path is bit-exact against the training code's own adapter path.
|
| 138 |
-
|
| 139 |
-
## Output sizes
|
| 140 |
-
|
| 141 |
-
Text-to-image and editing were distilled on different target buckets, so the size menu follows the
|
| 142 |
-
mode and rewrites itself as soon as a reference image is attached:
|
| 143 |
-
|
| 144 |
-
| mode | target areas | why |
|
| 145 |
-
|---|---|---|
|
| 146 |
-
| text-to-image (no references) | 1024² and 2048² | `build_user_manifest.py` draws T2I targets 50/50 from these |
|
| 147 |
-
| editing (1–6 references) | 1024² and 1536², plus *auto* | editing targets are drawn 70/30 from these |
|
| 148 |
-
|
| 149 |
-
Every entry is `calculate_dimensions(area, ratio)` rounded to a multiple of 32 — the pipeline's own
|
| 150 |
-
rule — for 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3. *Auto* takes the aspect ratio of the **last**
|
| 151 |
-
reference at 1024² area, which is what the pipeline does when `height`/`width` are left unset.
|
| 152 |
-
`generate()` additionally snaps any editing call whose size is not one of the offered editing
|
| 153 |
-
dimensions onto the 1536² bucket of the nearest aspect ratio, and says so in the status line, so an
|
| 154 |
-
API client or a stale dropdown cannot push a multi-reference call out of budget.
|
| 155 |
-
|
| 156 |
-
**Custom** (last entry of both menus) shows Width / Height sliders, 512 to 2720 in steps of 32, pre-filled
|
| 157 |
-
with the last preset picked. A custom size above **2048² area for text-to-image** or **1536² area for editing**
|
| 158 |
-
(the largest trained bucket of each mode) is scaled down to that area keeping its ratio, floored to multiples
|
| 159 |
-
of 32, and the status line says so; smaller sizes are used as given. Sizes off the trained buckets work but
|
| 160 |
-
were not evaluated.
|
| 161 |
-
|
| 162 |
-
Reference images are always resized to **1024² area** before encoding regardless of the output size,
|
| 163 |
-
so extra references cost VRAM and time but not output resolution.
|
| 164 |
|
| 165 |
## Prompt enhancement
|
| 166 |
|
| 167 |
-
**Enhance prompt**
|
| 168 |
-
|
| 169 |
-
|
| 170 |
-
|
| 171 |
-
|
| 172 |
-
Text-to-image prompts become one English scene paragraph of about 150 words that is told the chosen aspect ratio;
|
| 173 |
-
text to be rendered stays in the request's language. Editing instructions become one to three sentences in the
|
| 174 |
-
instruction's language that spell out what to keep, and the reference images go along, each downscaled to at
|
| 175 |
-
most **0.5 MP**. Measured from our cluster: median 2.0 s for text-to-image, 0.5 s for edits, about $0.0003 per call.
|
| 176 |
-
The request runs before, and outside, the `@spaces.GPU` function, so waiting on it holds no GPU and costs no
|
| 177 |
-
ZeroGPU quota. The rewritten prompt is shown under **Prompt sent to the model**; if the call fails (no
|
| 178 |
-
`OPENROUTER_API_KEY` secret, 30 s timeout, API error), the original prompt is used and the status line says so.
|
| 179 |
-
|
| 180 |
-
**Privacy.** When enhancement runs, the prompt and the downscaled reference images are sent to OpenRouter and the
|
| 181 |
-
third-party provider it routes the request to; the request only allows providers that neither train on nor retain
|
| 182 |
-
the data (OpenRouter's `data_collection: deny` and zero-data-retention filters). With **Off**, nothing leaves the Space.
|
| 183 |
-
|
| 184 |
-
Caveats. The student was distilled on raw user prompts, so enhancement is a convenience layered on top of the
|
| 185 |
-
recipe, not part of it. The rewrite is sampled, so the same prompt can come back different, and it sometimes adds
|
| 186 |
-
things the request did not ask for, on-image text included (in testing, one café sign gained an invented
|
| 187 |
-
opening-date line). Switch it **Off** when the wording matters (exact text to render, deliberately terse
|
| 188 |
-
prompts). Until this version the Space ran the same official system prompts on its own Qwen3-VL text encoder
|
| 189 |
-
(on the GPU, ~9 s per call); the three-reference example was rendered with that rewriter.
|
| 190 |
-
|
| 191 |
-
## Hardware
|
| 192 |
-
|
| 193 |
-
Weights are ~31.5 GiB resident in bf16 (13.25 transformer + 16.33 text encoder + 0.63 VAE + the r=256
|
| 194 |
-
adapter, 1.3 GiB, loaded in bf16). Measured on one B200 with the v0.2 LoRA at 5 steps (the 6-step default adds one transformer forward, about
|
| 195 |
-
+20% on the sampling time and no extra memory), with the earlier on-GPU prompt rewrite on (it ran before the
|
| 196 |
-
denoise and did not raise the peak), around a single `generate()` call including the VAE encode/decode, with the allocator
|
| 197 |
-
cache dropped before each measurement (`release/measure_mem.py`; the two 2048² rows are the v0.1 full-transformer
|
| 198 |
-
measurements at 4 steps plus the adapter). **No VAE tiling anywhere**: `vae.enable_tiling()` also tiles the
|
| 199 |
-
*encode* of the reference images (256 px tiles), and tiled reference latents wreck an edit - duplicated subjects,
|
| 200 |
-
wrong scale - while a tiled decode is not the validated pipeline either.
|
| 201 |
-
|
| 202 |
-
| call | output | peak allocated | peak reserved | wall clock (B200) |
|
| 203 |
-
|---|---|---|---|---|
|
| 204 |
-
| text-to-image | 1024×1024 | 38.24 GiB | 40.05 GiB | 0.8 s |
|
| 205 |
-
| text-to-image | 2048×2048 | ~58.3 GiB | ~66 GiB | 3.4 s (v0.1) |
|
| 206 |
-
| edit, 1 reference, auto | 1024×1024 | 40.13 GiB | 42.68 GiB | 1.1 s |
|
| 207 |
-
| edit, 2 references, auto | 832×1248 | 42.04 GiB | 44.71 GiB | 1.5 s |
|
| 208 |
-
| edit, 3 references, auto (the 3-reference example) | 928×1152 | 44.29 GiB | 47.84 GiB | 2.9 s |
|
| 209 |
-
| edit, 3 references (largest) | 1760×1344 | ~52.8 GiB | ~59 GiB | 2.9 s (v0.1) |
|
| 210 |
-
| edit, 6 references, auto | 1024×1024 | 50.17 GiB | 51.95 GiB | see below |
|
| 211 |
-
| edit, 6 references (largest, the `MAX_REFS` limit) | 1760×1344 | 58.41 GiB | 62.29 GiB | see below |
|
| 212 |
-
|
| 213 |
-
The two 6-reference rows are v0.2.1 at 6 steps on a B200 shared with other jobs, so their times are not
|
| 214 |
-
comparable with the rows above: warm, the denoise took 3.7 s at auto and 6.0 s at 1760×1344 while the node was
|
| 215 |
-
lightly loaded (9-22 s while it was busy), and the first call at a new output size 2-6 s longer. Each reference
|
| 216 |
-
is resized to 1024² area before encoding, so memory grows by ~2 GiB per reference whatever the upload size
|
| 217 |
-
(3 → 6 references: +5.9 GiB at auto, +6.0 GiB at 1760×1344).
|
| 218 |
-
|
| 219 |
-
The turbo was distilled on edits with at most three references (`build_user_manifest.py` drops samples with
|
| 220 |
-
more), so four to six are outside its training data. In one spot check with six distinct references (two women,
|
| 221 |
-
a cat, a bird, a pony and a shop front; seeds 1-3, enhancement on) all six appeared in every output, at auto and
|
| 222 |
-
at 1760×1344.
|
| 223 |
-
|
| 224 |
-
ZeroGPU `large` is a 48 GB card with 44.7 GiB usable, so the three-reference edit at the 1024² bucket is
|
| 225 |
-
exactly the call that failed there (the CUDA OOM surfaces on ZeroGPU as `NVML_SUCCESS == r INTERNAL ASSERT
|
| 226 |
-
FAILED` from the caching allocator, because the container cannot query NVML for the OOM report), and the
|
| 227 |
-
2048² text-to-image and 1536² editing buckets never fit untiled. `app.py` therefore requests
|
| 228 |
-
`@spaces.GPU(duration=90, size="xlarge")`: the full 96 GB RTX Pro 6000 Blackwell, at **2× quota** (an
|
| 229 |
-
unauthenticated visitor's 2 min/day covers about three of the largest calls, a free account's 5 min about
|
| 230 |
-
seven; PRO 40 min). `app.py` also sets `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`. Prompt enhancement
|
| 231 |
-
is an API call made before the `@spaces.GPU` function is entered, so it uses no GPU time. Wall-clock on the
|
| 232 |
-
ZeroGPU card will be slower than the B200 numbers above; `duration=90` covers the largest call.
|
| 233 |
-
|
| 234 |
-
**Cold start** is dominated by the download: ~30.9 GiB of base safetensors from
|
| 235 |
-
`Qwen/Qwen-Image-2.1` (13.25 transformer + 16.33 text encoder + 1.26 VAE + 0.02 processor) plus
|
| 236 |
-
**1.3 GiB** for the LoRA — the shipped root adapter is bf16, and `load_lora_weights(weight_name=…)`
|
| 237 |
-
fetches that single file, not the F32 `peft_v0.2.1/` copy — then ~30–40 s to load and pack the pipeline
|
| 238 |
-
onto the GPU. Only the adapter changes
|
| 239 |
-
between releases, so a cached Space restarts far faster than it first boots.
|
| 240 |
-
|
| 241 |
-
### Alpha channels
|
| 242 |
-
|
| 243 |
-
`gr.Image(..., image_mode="RGB")` drops any alpha channel on upload. The pipeline then calls
|
| 244 |
-
`img.convert("RGBA")` on every condition image itself — the Qwen-Image-2.1 VAE genuinely has
|
| 245 |
-
`in_channels=4` — so what it actually sees is your image with a **fully opaque** alpha. A
|
| 246 |
-
transparent PNG is therefore not composited onto a background: PIL's `RGBA→RGB` keeps the raw colour
|
| 247 |
-
of transparent pixels (often black), and those colours become visible content. Flatten transparent
|
| 248 |
-
images onto the background you want before uploading.
|
| 249 |
-
|
| 250 |
-
### The vision patch-embed workaround
|
| 251 |
|
| 252 |
-
|
| 253 |
-
|
| 254 |
-
|
| 255 |
-
pipeline (356 ms in isolation with memory to spare; the same conv in fp32 is 2.9 ms). The matmul is
|
| 256 |
-
**0.8 ms** and its output is bitwise equal to the conv's — all verification images are unchanged
|
| 257 |
-
byte-for-byte. Without it, a 1-reference edit takes 32.3 s and a 2-reference edit 63.1 s, which is why
|
| 258 |
-
`@spaces.GPU(duration=90)` is enough only with the fix in place. Remove it if transformers ever ships
|
| 259 |
-
a fast path here.
|
| 260 |
|
| 261 |
-
##
|
| 262 |
|
| 263 |
-
`
|
| 264 |
-
`
|
| 265 |
-
|
| 266 |
-
|
| 267 |
-
undo the transformers>=5.0 hidden-state tying; a different major may break it), plus
|
| 268 |
-
`accelerate==1.15.0`, `peft==0.18.0` and `spaces==0.51.3` — the versions this release was validated
|
| 269 |
-
against. `peft` is required: without it `load_lora_weights` has no runtime to attach the adapter to.
|
| 270 |
-
`torch` and `gradio` are **not** listed: `sdk_version` above owns the Gradio version, and the
|
| 271 |
-
ZeroGPU image provides torch (its docs list 2.8.0–2.13.0 as supported, which includes the 2.9.1 this
|
| 272 |
-
release was validated on). That last point is an assumption about the image, not a measurement —
|
| 273 |
-
check the first build log and pin `torch==2.9.1` if it resolves something outside that range.
|
| 274 |
|
| 275 |
## License
|
| 276 |
|
| 277 |
-
The base model is released under the
|
| 278 |
-
|
| 279 |
-
|
| 280 |
-
requires a separate license from Alibaba (`model-business@notice.qwencloud.com`). See `NOTICE` for
|
| 281 |
-
the attribution required by §3.c.
|
| 282 |
|
| 283 |
> Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi
|
| 284 |
> Laboratory Technology Co., Ltd. All Rights Reserved.
|
|
|
|
| 19 |
- text-to-image
|
| 20 |
- image-editing
|
| 21 |
- distillation
|
|
|
|
| 22 |
# Hardware: this Space needs ZeroGPU, set in Space Settings. `suggested_hardware` is deliberately
|
| 23 |
# unset, because the only ZeroGPU value the Hub metadata accepts is the legacy `zero-a10g`, and a
|
| 24 |
# 24 GB A10G cannot hold this pipeline. `app.py` requests the 96 GB card with
|
| 25 |
+
# @spaces.GPU(size="xlarge") (no VAE tiling).
|
| 26 |
# Secrets: set HF_TOKEN (read scope) while Viggle/Qwen-Image-2.1-viggle-turbo is private.
|
| 27 |
# `hf_oauth` is not needed: the app calls no user-scoped Hub API.
|
| 28 |
---
|
| 29 |
|
| 30 |
+
# Viggle Turbo v0.3 — 6-step Qwen-Image-2.1
|
| 31 |
|
| 32 |
+
A distilled [Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) that generates and edits images in
|
| 33 |
+
**6 steps** with **no classifier-free guidance**, about **5× faster** than the 40-step base model. On most prompts it
|
| 34 |
+
is hard to tell apart from the base model; small, dense text and complicated edits (multi-reference composition, face
|
| 35 |
+
swaps, identity-preserving edits) can still fall short of it.
|
|
|
|
| 36 |
|
| 37 |
+
**v0.3 (2026-09-29):** at 6 steps, less grain than v0.2.1 and a little softer on fine texture. We think 6 steps is
|
| 38 |
+
close to its capacity: every further gain we found cost something elsewhere. The new **9-step** setting runs 7 turbo
|
| 39 |
+
steps and lets the base model finish the last two: finer detail and small text right more often, at about 1.4–1.5× the
|
| 40 |
+
time of 6 steps.
|
| 41 |
+
Weights, numbers and ComfyUI workflows are on the
|
| 42 |
+
[model card](https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo).
|
| 43 |
|
| 44 |
+
**Generate:** leave **References** empty for text-to-image, or add one to six images for editing, composition or
|
| 45 |
+
style transfer; the prompt refers to them as image 1, image 2, … in the order they were added. **Steps** runs from 3
|
| 46 |
+
to 9, 6 is the default. The **Examples** rows are pre-rendered by this exact app; clicking one loads its prompt,
|
| 47 |
+
references, settings and seed. The reference photos `woman1/cat_window/bird.webp` and the three-reference prompt come
|
| 48 |
+
from the [black-forest-labs/flux-klein-9b-kv](https://huggingface.co/spaces/black-forest-labs/flux-klein-9b-kv) Space;
|
| 49 |
+
the `qwen_*` rows come from the [Qwen/Qwen-Image-2.1](https://huggingface.co/spaces/Qwen/Qwen-Image-2.1) Space.
|
| 50 |
|
| 51 |
+
**Comparison** (no GPU needed): 32 of the 37 examples of the Qwen/Qwen-Image-2.1 Space, pre-rendered with this model
|
| 52 |
+
and with the base model in 40 steps (same prompt, inputs and seed 42, prompt enhancement off, one sample each), in an
|
| 53 |
+
image slider, plus Qwen's own API output where their Space ships one.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 54 |
|
| 55 |
## Prompt enhancement
|
| 56 |
|
| 57 |
+
**Enhance prompt**: **Auto** (default) rewrites prompts shorter than 30 tokens, **On** always rewrites, **Off** never
|
| 58 |
+
does. The rewrite is done by DeepSeek V4.1 Flash through OpenRouter, prompted with shortened versions of the system
|
| 59 |
+
prompts published with `Qwen/Qwen-Image-2.1-PE-T2I` / `-PE-I2I`. The rewritten prompt is shown under **Prompt sent to
|
| 60 |
+
the model**; if the call fails, the original prompt is used. The rewrite can add things the request did not ask for,
|
| 61 |
+
on-image text included, so switch it **Off** when the wording matters.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 62 |
|
| 63 |
+
**Privacy.** When enhancement runs, the prompt and the reference images (downscaled to at most 0.5 MP) are sent to
|
| 64 |
+
OpenRouter and the provider it routes the request to; the request only allows providers that neither train on nor
|
| 65 |
+
retain the data. With **Off**, nothing leaves the Space.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 66 |
|
| 67 |
+
## Running it
|
| 68 |
|
| 69 |
+
Environment: `LORA_FILE` (default the v0.3 r256 adapter; the v0.2.1 and v0.2 adapters in the model repo also work),
|
| 70 |
+
`STUDENT_REPO`, `BASE_MODEL_ID`, `HF_TOKEN` (read scope, needed only while the model repo is private),
|
| 71 |
+
`OPENROUTER_API_KEY` (prompt enhancement; without it every prompt is sent as written). The app needs ZeroGPU `xlarge`
|
| 72 |
+
(96 GB): the largest calls do not fit on the 48 GB card untiled, and tiling the VAE breaks edits.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 73 |
|
| 74 |
## License
|
| 75 |
|
| 76 |
+
The base model is released under the [Qwen RESEARCH LICENSE AGREEMENT](./LICENSE) — **non-commercial: research or
|
| 77 |
+
evaluation purposes only**. This distilled derivative and this demo inherit that restriction. Commercial use requires
|
| 78 |
+
a separate license from Alibaba (`model-business@notice.qwencloud.com`). See `NOTICE` for the required attribution.
|
|
|
|
|
|
|
| 79 |
|
| 80 |
> Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi
|
| 81 |
> Laboratory Technology Co., Ltd. All Rights Reserved.
|
app.py
CHANGED
|
@@ -1,4 +1,4 @@
|
|
| 1 |
-
"""Gradio demo for the
|
| 2 |
|
| 3 |
# ZeroGPU patches torch at import time, so `spaces` must be imported before torch or any CUDA use.
|
| 4 |
# find_spec keeps this file runnable off-Spaces, where the package is absent.
|
|
@@ -34,20 +34,17 @@ import httpx
|
|
| 34 |
import torch
|
| 35 |
import compare
|
| 36 |
from PIL import Image, ImageOps
|
| 37 |
-
from diffusers import FlowMatchEulerDiscreteScheduler, QwenImage21Pipeline
|
| 38 |
from diffusers.pipelines.qwenimage21.pipeline_qwenimage21 import calculate_dimensions
|
| 39 |
from huggingface_hub import CommitScheduler
|
| 40 |
|
| 41 |
-
#
|
| 42 |
-
#
|
| 43 |
-
#
|
| 44 |
-
# v0.1 adapters are still in the repo. STUDENT="full": the v0.1 full fine-tune in transformer/ (bf16, exact, no
|
| 45 |
-
# adapter) replaces the base transformer at load time - the base transformer is then not downloaded at all.
|
| 46 |
BASE_MODEL_ID = os.environ.get("BASE_MODEL_ID", "Qwen/Qwen-Image-2.1")
|
| 47 |
STUDENT_REPO = os.environ.get("STUDENT_REPO", os.environ.get("LORA_REPO", "Viggle/Qwen-Image-2.1-viggle-turbo"))
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
STUDENT_TAG = {"full": "v0.1 full fine-tune (transformer/)", "lora": "v0.2.1 LoRA r256" if "v0.2.1" in LORA_FILE else f"LoRA {LORA_FILE}"}[STUDENT]
|
| 51 |
HF_TOKEN = os.environ.get("HF_TOKEN") # only needed while the model repo is private
|
| 52 |
# Usage counter. The Space autoscales to several replicas and its log view shows only one of them, so the [gen] log lines
|
| 53 |
# undercount. Every generate() call that reaches the GPU also appends one JSON line (metadata only: no prompt, no images)
|
|
@@ -81,7 +78,10 @@ def record(**row):
|
|
| 81 |
# launched with, 6 = default, 7 also validated; 4 = the training nodes; 3 is the pipeline's default linspace and
|
| 82 |
# off-contract). Passing sigmas= to the pipeline is the whole recipe - it applies the same resolution-dependent
|
| 83 |
# shift it applies to its default nodes.
|
|
|
|
|
|
|
| 84 |
STEPS = 6
|
|
|
|
| 85 |
|
| 86 |
|
| 87 |
def raw_nodes(steps):
|
|
@@ -90,6 +90,8 @@ def raw_nodes(steps):
|
|
| 90 |
return None # pipeline default linspace(1, 1/steps, steps); off-contract
|
| 91 |
if steps == 4:
|
| 92 |
return [1.0, 0.75, 0.5, 0.25]
|
|
|
|
|
|
|
| 93 |
return [1.0 - 0.125 * i / (steps - 4) for i in range(steps - 4)] + [0.875, 0.75, 0.5, 0.25]
|
| 94 |
|
| 95 |
|
|
@@ -104,8 +106,8 @@ EXAMPLES_DIR = Path(__file__).parent / "examples"
|
|
| 104 |
EXAMPLES = json.loads((EXAMPLES_DIR / "manifest.json").read_text(encoding="utf-8")) if (EXAMPLES_DIR / "manifest.json").exists() else []
|
| 105 |
|
| 106 |
# Output sizes are calculate_dimensions(area, ratio) — the pipeline's own rule, rounded to multiples
|
| 107 |
-
# of 32 — over the target areas the student was distilled on
|
| 108 |
-
#
|
| 109 |
# different menus. Reference images are always encoded at 1024^2 area regardless.
|
| 110 |
RATIOS = [("1:1", 1.0), ("16:9", 16 / 9), ("9:16", 9 / 16), ("4:3", 4 / 3), ("3:4", 3 / 4), ("3:2", 3 / 2), ("2:3", 2 / 3)]
|
| 111 |
AUTO = "Auto · match the last reference (1024² area)"
|
|
@@ -131,18 +133,11 @@ EDIT_DIMS = [SIZES[label] for label in EDIT_CHOICES if SIZES[label]]
|
|
| 131 |
CUSTOM_CAP = {False: 2048, True: 1536} # keyed by "has references"
|
| 132 |
MAX_SIDE = max(max(size) for size in SIZES.values() if size)
|
| 133 |
|
| 134 |
-
|
| 135 |
-
|
| 136 |
-
STUDENT_REPO, subfolder="transformer", torch_dtype=torch.bfloat16, token=HF_TOKEN
|
| 137 |
-
)
|
| 138 |
-
pipe = QwenImage21Pipeline.from_pretrained(BASE_MODEL_ID, transformer=transformer, dtype=torch.bfloat16, token=HF_TOKEN)
|
| 139 |
-
else:
|
| 140 |
-
pipe = QwenImage21Pipeline.from_pretrained(BASE_MODEL_ID, dtype=torch.bfloat16, token=HF_TOKEN)
|
| 141 |
-
pipe.load_lora_weights(STUDENT_REPO, weight_name=LORA_FILE, token=HF_TOKEN)
|
| 142 |
# The stock Qwen-Image-2.1 scheduler config carries shift_terminal=0.02, which stretches the last
|
| 143 |
# sigma node to 0.02 instead of 0 and silently costs the final step. The student was distilled
|
| 144 |
-
# against the unstretched schedule, so rebuild the scheduler with shift_terminal=None
|
| 145 |
-
# pipe(num_inference_steps=4) alone reproduces dmd_common.student_sigmas(terminal="none") exactly.
|
| 146 |
pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_config(pipe.scheduler.config, shift_terminal=None)
|
| 147 |
pipe.to("cuda")
|
| 148 |
|
|
@@ -214,12 +209,37 @@ def enhance_prompt(prompt, images, ratio_name):
|
|
| 214 |
return None
|
| 215 |
|
| 216 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 217 |
@gpu
|
| 218 |
def denoise(prompt, images, width, height, steps, seed):
|
| 219 |
# No VAE tiling anywhere (xlarge card): tiling the reference encode wrecks edits, and tiled decodes are not the
|
| 220 |
# validated pipeline either.
|
| 221 |
generator = torch.Generator(device="cuda").manual_seed(seed)
|
| 222 |
started = time.perf_counter()
|
|
|
|
|
|
|
| 223 |
result = pipe(
|
| 224 |
prompt=prompt,
|
| 225 |
image=images or None,
|
|
@@ -230,13 +250,15 @@ def denoise(prompt, images, width, height, steps, seed):
|
|
| 230 |
true_cfg_scale=1.0, # the student is distilled without classifier-free guidance; keep it off
|
| 231 |
output_resolution=1024, # condition images are encoded at 1024-area, as in distillation
|
| 232 |
generator=generator,
|
|
|
|
| 233 |
).images[0]
|
|
|
|
| 234 |
return result, time.perf_counter() - started
|
| 235 |
|
| 236 |
|
| 237 |
def generate(prompt, refs, size_label, seed, randomize_seed, enhance="Auto", steps=STEPS, custom_width=1024, custom_height=1024):
|
| 238 |
t0 = time.perf_counter()
|
| 239 |
-
steps = min(max(int(steps), 3),
|
| 240 |
# refs: the gallery's (filepath, caption) pairs in upload order, which is the order the prompt's "image 1, 2, ..." refer
|
| 241 |
# to. EXIF rotation + RGB is what the gr.Image slots used to do.
|
| 242 |
if len(refs or []) > MAX_REFS:
|
|
@@ -357,9 +379,9 @@ REFS_CSS = """
|
|
| 357 |
#add-ref { display: none; }
|
| 358 |
""" + f"body:has(#refs .gallery-item):not(:has(#refs .gallery-item:nth-child({MAX_REFS}))) #add-ref {{ display: flex; }}"
|
| 359 |
|
| 360 |
-
with gr.Blocks(title="Viggle Turbo v0.
|
| 361 |
gr.Markdown(
|
| 362 |
-
"# Viggle Turbo v0.
|
| 363 |
"Text-to-image and image editing in **6 steps**: about **5× faster** than the 40-step Qwen-Image-2.1 and very "
|
| 364 |
"competitive with it in quality — see the **Comparison** tab. Leave the references empty for text-to-image, "
|
| 365 |
f"or add up to {MAX_REFS} to edit, compose or transfer style.\n\n"
|
|
@@ -368,12 +390,14 @@ with gr.Blocks(title="Viggle Turbo v0.2.1 · Qwen-Image-2.1 6-step", theme=THEME
|
|
| 368 |
)
|
| 369 |
with gr.Accordion("About this model", open=False):
|
| 370 |
gr.Markdown(
|
| 371 |
-
"A
|
| 372 |
-
"
|
| 373 |
-
"
|
| 374 |
"Qwen Space side by side, turbo vs base, same seed.\n\n"
|
| 375 |
-
"**v0.
|
| 376 |
-
"
|
|
|
|
|
|
|
| 377 |
f"Weights: **{STUDENT_TAG}** from [{STUDENT_REPO}](https://huggingface.co/{STUDENT_REPO})."
|
| 378 |
)
|
| 379 |
with gr.Tabs() as tabs:
|
|
@@ -406,8 +430,8 @@ with gr.Blocks(title="Viggle Turbo v0.2.1 · Qwen-Image-2.1 6-step", theme=THEME
|
|
| 406 |
"sent to the image model is shown under the result.",
|
| 407 |
)
|
| 408 |
with gr.Accordion("Advanced settings", open=False) as advanced:
|
| 409 |
-
steps = gr.Slider(label="Steps", minimum=3, maximum=
|
| 410 |
-
info="6 is the
|
| 411 |
with gr.Row():
|
| 412 |
seed = gr.Number(label="Seed", value=0, precision=0, interactive=False)
|
| 413 |
randomize_seed = gr.Checkbox(label="Randomize seed", value=True)
|
|
|
|
| 1 |
+
"""Gradio demo for the distilled Qwen-Image-2.1 turbo (v0.3: LoRA r256 sampled in 6 steps; 9 steps = 7 turbo + 2 base-model steps)."""
|
| 2 |
|
| 3 |
# ZeroGPU patches torch at import time, so `spaces` must be imported before torch or any CUDA use.
|
| 4 |
# find_spec keeps this file runnable off-Spaces, where the package is absent.
|
|
|
|
| 34 |
import torch
|
| 35 |
import compare
|
| 36 |
from PIL import Image, ImageOps
|
| 37 |
+
from diffusers import FlowMatchEulerDiscreteScheduler, QwenImage21Pipeline
|
| 38 |
from diffusers.pipelines.qwenimage21.pipeline_qwenimage21 import calculate_dimensions
|
| 39 |
from huggingface_hub import CommitScheduler
|
| 40 |
|
| 41 |
+
# The v0.3 LoRA (r256) at the model repo root, applied at runtime on top of the base transformer and never merged
|
| 42 |
+
# (merging into bf16 keeps only ~47% of the adapter delta, the per-element deltas sit below the bf16 ULP of the base
|
| 43 |
+
# weights). LORA_FILE selects the adapter file; the v0.2.1 and v0.2 adapters are still in the repo.
|
|
|
|
|
|
|
| 44 |
BASE_MODEL_ID = os.environ.get("BASE_MODEL_ID", "Qwen/Qwen-Image-2.1")
|
| 45 |
STUDENT_REPO = os.environ.get("STUDENT_REPO", os.environ.get("LORA_REPO", "Viggle/Qwen-Image-2.1-viggle-turbo"))
|
| 46 |
+
LORA_FILE = os.environ.get("LORA_FILE", "Qwen-Image-2.1-viggle-turbo-v0.3-6step-lora-r256.safetensors")
|
| 47 |
+
STUDENT_TAG = "v0.3 LoRA r256" if "v0.3" in LORA_FILE else f"LoRA {LORA_FILE}"
|
|
|
|
| 48 |
HF_TOKEN = os.environ.get("HF_TOKEN") # only needed while the model repo is private
|
| 49 |
# Usage counter. The Space autoscales to several replicas and its log view shows only one of them, so the [gen] log lines
|
| 50 |
# undercount. Every generate() call that reaches the GPU also appends one JSON line (metadata only: no prompt, no images)
|
|
|
|
| 78 |
# launched with, 6 = default, 7 also validated; 4 = the training nodes; 3 is the pipeline's default linspace and
|
| 79 |
# off-contract). Passing sigmas= to the pipeline is the whole recipe - it applies the same resolution-dependent
|
| 80 |
# shift it applies to its default nodes.
|
| 81 |
+
# 9 is different: the 7-step schedule, one more student step 0.25 -> 1/6, then the base model (LoRA disabled) for the
|
| 82 |
+
# last two steps 1/6 -> 1/12 -> 0 (BASE_FROM = first base-model step).
|
| 83 |
STEPS = 6
|
| 84 |
+
BASE_FROM = 7
|
| 85 |
|
| 86 |
|
| 87 |
def raw_nodes(steps):
|
|
|
|
| 90 |
return None # pipeline default linspace(1, 1/steps, steps); off-contract
|
| 91 |
if steps == 4:
|
| 92 |
return [1.0, 0.75, 0.5, 0.25]
|
| 93 |
+
if steps == 9:
|
| 94 |
+
return raw_nodes(7) + [1 / 6, 1 / 12]
|
| 95 |
return [1.0 - 0.125 * i / (steps - 4) for i in range(steps - 4)] + [0.875, 0.75, 0.5, 0.25]
|
| 96 |
|
| 97 |
|
|
|
|
| 106 |
EXAMPLES = json.loads((EXAMPLES_DIR / "manifest.json").read_text(encoding="utf-8")) if (EXAMPLES_DIR / "manifest.json").exists() else []
|
| 107 |
|
| 108 |
# Output sizes are calculate_dimensions(area, ratio) — the pipeline's own rule, rounded to multiples
|
| 109 |
+
# of 32 — over the target areas the student was distilled on: 1024^2 / 2048^2 for text-to-image and
|
| 110 |
+
# 1024^2 / 1536^2 for editing, so the two modes get
|
| 111 |
# different menus. Reference images are always encoded at 1024^2 area regardless.
|
| 112 |
RATIOS = [("1:1", 1.0), ("16:9", 16 / 9), ("9:16", 9 / 16), ("4:3", 4 / 3), ("3:4", 3 / 4), ("3:2", 3 / 2), ("2:3", 2 / 3)]
|
| 113 |
AUTO = "Auto · match the last reference (1024² area)"
|
|
|
|
| 133 |
CUSTOM_CAP = {False: 2048, True: 1536} # keyed by "has references"
|
| 134 |
MAX_SIDE = max(max(size) for size in SIZES.values() if size)
|
| 135 |
|
| 136 |
+
pipe = QwenImage21Pipeline.from_pretrained(BASE_MODEL_ID, dtype=torch.bfloat16, token=HF_TOKEN)
|
| 137 |
+
pipe.load_lora_weights(STUDENT_REPO, weight_name=LORA_FILE, token=HF_TOKEN)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 138 |
# The stock Qwen-Image-2.1 scheduler config carries shift_terminal=0.02, which stretches the last
|
| 139 |
# sigma node to 0.02 instead of 0 and silently costs the final step. The student was distilled
|
| 140 |
+
# against the unstretched schedule, so rebuild the scheduler with shift_terminal=None.
|
|
|
|
| 141 |
pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_config(pipe.scheduler.config, shift_terminal=None)
|
| 142 |
pipe.to("cuda")
|
| 143 |
|
|
|
|
| 209 |
return None
|
| 210 |
|
| 211 |
|
| 212 |
+
# The pipeline computes the text/reference K/V once at step 0 ("extract") and reuses them ("cached"). In 9-step mode
|
| 213 |
+
# those K/V come from the student, so the first base-model step extracts them again with the LoRA off. The output is
|
| 214 |
+
# bitwise equal to running without the cache, at the cost of one extra extract instead of one per step.
|
| 215 |
+
_transformer_forward = pipe.transformer.forward
|
| 216 |
+
reextract = [False]
|
| 217 |
+
|
| 218 |
+
|
| 219 |
+
def transformer_forward(*args, kv_cache_mode=None, **kwargs):
|
| 220 |
+
if reextract[0] and kv_cache_mode == "cached":
|
| 221 |
+
kv_cache_mode, reextract[0] = "extract", False
|
| 222 |
+
return _transformer_forward(*args, kv_cache_mode=kv_cache_mode, **kwargs)
|
| 223 |
+
|
| 224 |
+
|
| 225 |
+
pipe.transformer.forward = transformer_forward
|
| 226 |
+
|
| 227 |
+
|
| 228 |
+
def base_tail(pipe, i, t, kwargs):
|
| 229 |
+
if i == BASE_FROM - 1: # runs after step i, so steps BASE_FROM.. are the base model's
|
| 230 |
+
pipe.disable_lora()
|
| 231 |
+
reextract[0] = True
|
| 232 |
+
return kwargs
|
| 233 |
+
|
| 234 |
+
|
| 235 |
@gpu
|
| 236 |
def denoise(prompt, images, width, height, steps, seed):
|
| 237 |
# No VAE tiling anywhere (xlarge card): tiling the reference encode wrecks edits, and tiled decodes are not the
|
| 238 |
# validated pipeline either.
|
| 239 |
generator = torch.Generator(device="cuda").manual_seed(seed)
|
| 240 |
started = time.perf_counter()
|
| 241 |
+
pipe.enable_lora() # every call starts on the student, even if an earlier 9-step call died with the LoRA off
|
| 242 |
+
reextract[0] = False
|
| 243 |
result = pipe(
|
| 244 |
prompt=prompt,
|
| 245 |
image=images or None,
|
|
|
|
| 250 |
true_cfg_scale=1.0, # the student is distilled without classifier-free guidance; keep it off
|
| 251 |
output_resolution=1024, # condition images are encoded at 1024-area, as in distillation
|
| 252 |
generator=generator,
|
| 253 |
+
callback_on_step_end=base_tail if steps == 9 else None,
|
| 254 |
).images[0]
|
| 255 |
+
pipe.enable_lora()
|
| 256 |
return result, time.perf_counter() - started
|
| 257 |
|
| 258 |
|
| 259 |
def generate(prompt, refs, size_label, seed, randomize_seed, enhance="Auto", steps=STEPS, custom_width=1024, custom_height=1024):
|
| 260 |
t0 = time.perf_counter()
|
| 261 |
+
steps = min(max(int(steps), 3), 9) # 6 (RAW_NODES) is the default; 9 = 7 student steps + 2 base-model steps; fewer than 6 degrade quickly
|
| 262 |
# refs: the gallery's (filepath, caption) pairs in upload order, which is the order the prompt's "image 1, 2, ..." refer
|
| 263 |
# to. EXIF rotation + RGB is what the gr.Image slots used to do.
|
| 264 |
if len(refs or []) > MAX_REFS:
|
|
|
|
| 379 |
#add-ref { display: none; }
|
| 380 |
""" + f"body:has(#refs .gallery-item):not(:has(#refs .gallery-item:nth-child({MAX_REFS}))) #add-ref {{ display: flex; }}"
|
| 381 |
|
| 382 |
+
with gr.Blocks(title="Viggle Turbo v0.3 · Qwen-Image-2.1 6-step", theme=THEME, head=FONT_HEAD + REFS_HEAD, css=REFS_CSS) as demo:
|
| 383 |
gr.Markdown(
|
| 384 |
+
"# Viggle Turbo v0.3 — 6-step Qwen-Image-2.1\n"
|
| 385 |
"Text-to-image and image editing in **6 steps**: about **5× faster** than the 40-step Qwen-Image-2.1 and very "
|
| 386 |
"competitive with it in quality — see the **Comparison** tab. Leave the references empty for text-to-image, "
|
| 387 |
f"or add up to {MAX_REFS} to edit, compose or transfer style.\n\n"
|
|
|
|
| 390 |
)
|
| 391 |
with gr.Accordion("About this model", open=False):
|
| 392 |
gr.Markdown(
|
| 393 |
+
"A distilled Qwen-Image-2.1 that runs with no classifier-free guidance. On most prompts it is hard to tell "
|
| 394 |
+
"apart from the base model; small, dense text and complicated edits (multi-reference composition, face swaps, "
|
| 395 |
+
"identity-preserving edits) can still fall short of it. The **Comparison** tab has 32 examples of the official "
|
| 396 |
"Qwen Space side by side, turbo vs base, same seed.\n\n"
|
| 397 |
+
"**v0.3 (2026-09-29):** at 6 steps, less grain than v0.2.1 and a little softer on fine texture. We think 6 steps "
|
| 398 |
+
"is close to its capacity: every further gain we found cost something elsewhere. **9 steps** runs 7 turbo "
|
| 399 |
+
"steps and lets the base model finish the last two: finer detail and small text right more often, at about 1.4–1.5× the "
|
| 400 |
+
"time of 6 steps. "
|
| 401 |
f"Weights: **{STUDENT_TAG}** from [{STUDENT_REPO}](https://huggingface.co/{STUDENT_REPO})."
|
| 402 |
)
|
| 403 |
with gr.Tabs() as tabs:
|
|
|
|
| 430 |
"sent to the image model is shown under the result.",
|
| 431 |
)
|
| 432 |
with gr.Accordion("Advanced settings", open=False) as advanced:
|
| 433 |
+
steps = gr.Slider(label="Steps", minimum=3, maximum=9, step=1, value=STEPS,
|
| 434 |
+
info="6 is the default; 7-8 add steps at the high-noise end; 9 = 7 turbo + 2 base-model steps (finer detail; ~1.4–1.5× the time).")
|
| 435 |
with gr.Row():
|
| 436 |
seed = gr.Number(label="Seed", value=0, precision=0, interactive=False)
|
| 437 |
randomize_seed = gr.Checkbox(label="Randomize seed", value=True)
|
compare/cases.json
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|
compare/img/case00/turbo.webp
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
compare/img/case00/turbo9.webp
ADDED
|
Git LFS Details
|
compare/img/{case01/turbo8.webp → case00/turbo9_t.webp}
RENAMED
|
File without changes
|
compare/img/case00/turbo_t.webp
CHANGED
|
|
Git LFS Details
|
compare/img/case01/turbo.webp
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
compare/img/case01/turbo8_t.webp
DELETED
|
Binary file (80.3 kB)
|
|
|
compare/img/{case17/turbo8.webp → case01/turbo9.webp}
RENAMED
|
File without changes
|
compare/img/case01/turbo9_t.webp
ADDED
|
compare/img/case01/turbo_t.webp
CHANGED
|
|
compare/img/case02/turbo.webp
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
compare/img/{case16/turbo8.webp → case02/turbo9.webp}
RENAMED
|
File without changes
|
compare/img/case02/turbo9_t.webp
ADDED
|
compare/img/case02/turbo_t.webp
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
compare/img/case03/turbo.webp
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
compare/img/{case18/turbo8.webp → case03/turbo9.webp}
RENAMED
|
File without changes
|
compare/img/case03/turbo9_t.webp
ADDED
|
Git LFS Details
|
compare/img/case03/turbo_t.webp
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
compare/img/case04/turbo.webp
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
compare/img/case04/turbo9.webp
ADDED
|
Git LFS Details
|
compare/img/case04/turbo9_t.webp
ADDED
|
compare/img/case04/turbo_t.webp
CHANGED
|
|
compare/img/case05/turbo.webp
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
compare/img/case05/turbo9.webp
ADDED
|
Git LFS Details
|
compare/img/case05/turbo9_t.webp
ADDED
|
compare/img/case05/turbo_t.webp
CHANGED
|
|
compare/img/case06/turbo.webp
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
compare/img/case06/turbo9.webp
ADDED
|
Git LFS Details
|
compare/img/case06/turbo9_t.webp
ADDED
|
compare/img/case06/turbo_t.webp
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
compare/img/case07/turbo.webp
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
compare/img/case07/turbo9.webp
ADDED
|
Git LFS Details
|
compare/img/case07/turbo9_t.webp
ADDED
|
compare/img/case07/turbo_t.webp
CHANGED
|
|
compare/img/case08/turbo.webp
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
compare/img/case08/turbo9.webp
ADDED
|
Git LFS Details
|
compare/img/case08/turbo9_t.webp
ADDED
|
compare/img/case08/turbo_t.webp
CHANGED
|
|
compare/img/case09/turbo.webp
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
compare/img/case09/turbo9.webp
ADDED
|
Git LFS Details
|
compare/img/case09/turbo9_t.webp
ADDED
|
compare/img/case09/turbo_t.webp
CHANGED
|
|
compare/img/case10/turbo.webp
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
compare/img/case10/turbo9.webp
ADDED
|
Git LFS Details
|
compare/img/case10/turbo9_t.webp
ADDED
|
compare/img/case10/turbo_t.webp
CHANGED
|
|