yycc commited on
Commit
4825513
·
verified ·
1 Parent(s): a8e67a7

Fix example clicks (cache_examples=False); Custom output size (2048² T2I / 1536² edit cap); UI: short header, Advanced settings, examples reproduce with their seed, Viggle theme

Browse files
Files changed (2) hide show
  1. README.md +8 -1
  2. app.py +86 -35
README.md CHANGED
@@ -56,7 +56,8 @@ app produced for it (fixed seed; 6 steps and prompt enhancement on unless the ro
56
  otherwise — the launch-graphic row is 8 steps at the 3:2 2048² size with the 3k-character brief sent as
57
  written, and the five `qwen_*` rows are sent as written at 2048² / 1536²), so you can see what the model
58
  does without spending GPU time — click a row to load the prompt, the references, the result, and the
59
- steps / size / enhancement settings that produced it. `release/render_examples.py` regenerates the rows
 
60
  and `examples/manifest.json` whenever the weights change. The reference photos
61
  `woman1/woman2/cat/cat_window/bird.webp` and the three-reference prompt come from the
62
  [black-forest-labs/flux-klein-9b-kv](https://huggingface.co/spaces/black-forest-labs/flux-klein-9b-kv) Space;
@@ -146,6 +147,12 @@ reference at 1024² area, which is what the pipeline does when `height`/`width`
146
  dimensions onto the 1536² bucket of the nearest aspect ratio, and says so in the status line, so an
147
  API client or a stale dropdown cannot push a three-reference call out of budget.
148
 
 
 
 
 
 
 
149
  Reference images are always resized to **1024² area** before encoding regardless of the output size,
150
  so extra references cost VRAM and time but not output resolution.
151
 
 
56
  otherwise — the launch-graphic row is 8 steps at the 3:2 2048² size with the 3k-character brief sent as
57
  written, and the five `qwen_*` rows are sent as written at 2048² / 1536²), so you can see what the model
58
  does without spending GPU time — click a row to load the prompt, the references, the result, and the
59
+ steps / size / enhancement settings and seed that produced it (with *Randomize seed* switched off), so
60
+ **Generate** reproduces the row. `release/render_examples.py` regenerates the rows
61
  and `examples/manifest.json` whenever the weights change. The reference photos
62
  `woman1/woman2/cat/cat_window/bird.webp` and the three-reference prompt come from the
63
  [black-forest-labs/flux-klein-9b-kv](https://huggingface.co/spaces/black-forest-labs/flux-klein-9b-kv) Space;
 
147
  dimensions onto the 1536² bucket of the nearest aspect ratio, and says so in the status line, so an
148
  API client or a stale dropdown cannot push a three-reference call out of budget.
149
 
150
+ **Custom** (last entry of both menus) shows Width / Height sliders, 512 to 2720 in steps of 32, pre-filled
151
+ with the last preset picked. A custom size above **2048² area for text-to-image** or **1536² area for editing**
152
+ (the largest trained bucket of each mode) is scaled down to that area keeping its ratio, floored to multiples
153
+ of 32, and the status line says so; smaller sizes are used as given. Sizes off the trained buckets work but
154
+ were not evaluated.
155
+
156
  Reference images are always resized to **1024² area** before encoding regardless of the output size,
157
  so extra references cost VRAM and time but not output resolution.
158
 
app.py CHANGED
@@ -78,6 +78,7 @@ EXAMPLES = json.loads((EXAMPLES_DIR / "manifest.json").read_text(encoding="utf-8
78
  # different menus. Reference images are always encoded at 1024^2 area regardless.
79
  RATIOS = [("1:1", 1.0), ("16:9", 16 / 9), ("9:16", 9 / 16), ("4:3", 4 / 3), ("3:4", 3 / 4), ("3:2", 3 / 2), ("2:3", 2 / 3)]
80
  AUTO = "Auto · match the last reference (1024² area)"
 
81
 
82
 
83
  def _bucket(area):
@@ -88,11 +89,16 @@ def _bucket(area):
88
  return sizes
89
 
90
 
91
- SIZES = {AUTO: None, **_bucket(1024), **_bucket(1536), **_bucket(2048)}
92
- T2I_CHOICES = [*_bucket(1024), *_bucket(2048)]
93
- EDIT_CHOICES = [AUTO, *_bucket(1024), *_bucket(1536)]
94
  # the (width, height) pairs the editing menu actually offers, in menu order
95
  EDIT_DIMS = [SIZES[label] for label in EDIT_CHOICES if SIZES[label]]
 
 
 
 
 
96
 
97
  if STUDENT == "full":
98
  transformer = QwenImage21Transformer2DModel.from_pretrained(
@@ -168,27 +174,36 @@ def enhance_prompt(prompt, images, ratio_name):
168
 
169
 
170
  @gpu
171
- def generate(prompt, image_1, image_2, image_3, size_label, seed, randomize_seed, enhance=True, steps=STEPS):
 
172
  steps = min(max(int(steps), 3), 8) # 6 (RAW_NODES) is the validated schedule; 3-8 are accepted, fewer or more degrade quickly
173
  images = [image for image in (image_1, image_2, image_3) if image is not None]
174
  # unchecked "Randomize seed" always honours the box, so the UI never says one thing and does another
175
  seed = random.randint(0, MAX_SEED) if (randomize_seed or seed is None) else max(int(seed), 0)
176
  size = SIZES.get(size_label)
177
  width, height = size if size else (None, None)
 
 
 
 
 
 
 
 
 
178
  # Belt and braces: an editing call must land on one of the editing buckets even if the caller
179
  # reaches generate() with a text-to-image size (stale dropdown, gr.Examples, API client). Test
180
  # membership in the offered dims rather than a raw area threshold: calculate_dimensions(1536**2,
181
  # 4/3) is 1760x1344 = 2_365_440 px, slightly *above* 1536**2, so an area test would flag two
182
  # sizes the editing menu itself offers. Snap to the 1536² bucket of the nearest named ratio.
183
- clamped = ""
184
- if images and width and (width, height) not in EDIT_DIMS:
185
  ratio = min(RATIOS, key=lambda name_ratio: abs(name_ratio[1] - width / height))[1]
186
  width, height, _ = calculate_dimensions(1536 * 1536, ratio)
187
  clamped = f" · clamped to {width}×{height} (editing caps at the 1536² bucket)"
188
  used_prompt, enhance_note = prompt, ""
189
  if enhance and prompt.strip():
190
  started = time.perf_counter()
191
- used_prompt = enhance_prompt(prompt, images, size_label.split(" · ")[0])
192
  enhance_note = f" · enhance {time.perf_counter() - started:.1f}s"
193
  if used_prompt == prompt:
194
  enhance_note += " (rewrite failed to parse, original prompt used)"
@@ -217,20 +232,44 @@ def refresh_sizes(image_1, image_2, image_3, current):
217
  return gr.update(choices=choices, value=current if current in choices else choices[0])
218
 
219
 
220
- with gr.Blocks(title="Viggle Turbo v0.2.1 · Qwen-Image-2.1 6-step") as demo:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
221
  gr.Markdown(
222
  "# Viggle Turbo v0.2.1 — 6-step Qwen-Image-2.1\n"
223
- "A DMD-distilled student of **Qwen-Image-2.1** that generates and edits in **6 sampling steps**, with no "
224
- "classifier-free guidance: about **5× faster** than the 40-step base model and very competitive with it in quality, "
225
- "and on some prompts we prefer its output. The **Comparison** tab has 33 examples of the official Qwen Space side by "
226
- "side, turbo vs base, same seed. Complicated edits (multi-reference composition, face swaps, identity-preserving "
227
- "edits) can still fall short of the base model.\n\n"
228
- "Leave the reference images empty for text-to-image; add one to three of them to edit, compose or transfer style. "
229
- "The size menu switches to the editing buckets as soon as a reference is attached. "
230
- "**v0.2.1 (2026-09-24):** the step-700 checkpoint of the v0.2 run on the 6-step schedule — intra-prompt diversity "
231
- "0.98× the base model, 0% composition drift; earlier versions and the numbers are on the model card. "
232
- f"Weights: **{STUDENT_TAG}** from [{STUDENT_REPO}](https://huggingface.co/{STUDENT_REPO})."
233
  )
 
 
 
 
 
 
 
 
 
 
234
  with gr.Tab("Generate"):
235
  with gr.Row():
236
  with gr.Column(scale=3):
@@ -239,23 +278,28 @@ with gr.Blocks(title="Viggle Turbo v0.2.1 · Qwen-Image-2.1 6-step") as demo:
239
  image_1 = gr.Image(label="Reference 1 (optional)", type="pil", image_mode="RGB", height=200)
240
  image_2 = gr.Image(label="Reference 2 (optional)", type="pil", image_mode="RGB", height=200)
241
  image_3 = gr.Image(label="Reference 3 (optional)", type="pil", image_mode="RGB", height=200)
242
- with gr.Row():
243
- # allow_custom_value: gradio validates API calls against the *initial* choices (the text-to-image
244
- # menu), which would reject every editing size sent through /generate; generate() snaps
245
- # anything off-menu itself.
246
- size_label = gr.Dropdown(label="Output size", choices=T2I_CHOICES, value=T2I_CHOICES[0], allow_custom_value=True)
247
- steps = gr.Slider(label="Steps", minimum=3, maximum=8, step=1, value=STEPS,
248
- info="6 is the validated schedule. Extra steps are always added at the high-noise end (7 also validated, 8 untested); fewer steps drop them again (4 = training nodes, 3 = uniform spacing, off-contract).")
249
- with gr.Row():
250
- seed = gr.Number(label="Seed", value=0, precision=0, interactive=False)
251
- randomize_seed = gr.Checkbox(label="Randomize seed", value=True)
252
  enhance = gr.Checkbox(
253
  label="Enhance prompt",
254
  value=True,
255
  info="Rewrite the prompt with the official Qwen-Image prompt-enhancement instructions "
256
  "(runs on the built-in Qwen3-VL text encoder, +4–15 s). The text actually sent is shown under the result.",
257
  )
258
- run = gr.Button("Generate", variant="primary")
 
 
 
 
 
 
259
  with gr.Column(scale=4):
260
  output_image = gr.Image(label="Result", type="pil", height=560)
261
  info = gr.Markdown()
@@ -263,11 +307,13 @@ with gr.Blocks(title="Viggle Turbo v0.2.1 · Qwen-Image-2.1 6-step") as demo:
263
 
264
  def example_details(prompt, image_1, image_2, image_3, result):
265
  """Runs on an example click (no GPU): fills the status line and the prompt stored with the row, and sets the
266
- steps / size / enhancement controls to what produced it (rows differ; Run then reproduces the result). It also sets
267
- the size menu's mode itself: loading an example does not fire the reference images' .input listeners."""
 
268
  row = next(row for row in EXAMPLES if row["prompt"] == prompt)
269
  choices = EDIT_CHOICES if any(row["refs"]) else T2I_CHOICES
270
- return row["info"], row["used_prompt"], row["steps"], gr.update(choices=choices, value=row["size_label"]), row["enhance"]
 
271
 
272
  if EXAMPLES:
273
  gr.Examples(
@@ -277,10 +323,13 @@ with gr.Blocks(title="Viggle Turbo v0.2.1 · Qwen-Image-2.1 6-step") as demo:
277
  ],
278
  inputs=[prompt, image_1, image_2, image_3, output_image],
279
  fn=example_details,
280
- outputs=[info, used_prompt, steps, size_label, enhance],
281
  run_on_click=True,
 
 
 
282
  examples_per_page=12,
283
- label="Examples · results pre-rendered by this model — click a row to load it (6 steps, prompt enhancement on, unless the status line says otherwise)",
284
  )
285
 
286
  with gr.Tab("Comparison"):
@@ -303,10 +352,12 @@ with gr.Blocks(title="Viggle Turbo v0.2.1 · Qwen-Image-2.1 6-step") as demo:
303
  inputs=[image_1, image_2, image_3, size_label],
304
  outputs=size_label,
305
  )
 
 
306
  randomize_seed.change(lambda on: gr.update(interactive=not on), inputs=randomize_seed, outputs=seed)
307
  run.click(
308
  generate,
309
- inputs=[prompt, image_1, image_2, image_3, size_label, seed, randomize_seed, enhance, steps],
310
  outputs=[output_image, info, used_prompt],
311
  )
312
 
 
78
  # different menus. Reference images are always encoded at 1024^2 area regardless.
79
  RATIOS = [("1:1", 1.0), ("16:9", 16 / 9), ("9:16", 9 / 16), ("4:3", 4 / 3), ("3:4", 3 / 4), ("3:2", 3 / 2), ("2:3", 2 / 3)]
80
  AUTO = "Auto · match the last reference (1024² area)"
81
+ CUSTOM = "Custom · set width and height below"
82
 
83
 
84
  def _bucket(area):
 
89
  return sizes
90
 
91
 
92
+ SIZES = {AUTO: None, CUSTOM: None, **_bucket(1024), **_bucket(1536), **_bucket(2048)}
93
+ T2I_CHOICES = [*_bucket(1024), *_bucket(2048), CUSTOM]
94
+ EDIT_CHOICES = [AUTO, *_bucket(1024), *_bucket(1536), CUSTOM]
95
  # the (width, height) pairs the editing menu actually offers, in menu order
96
  EDIT_DIMS = [SIZES[label] for label in EDIT_CHOICES if SIZES[label]]
97
+ # Custom sizes keep their aspect ratio but are scaled into the mode's largest distilled bucket, which is also what the
98
+ # memory table in the README measured: 2048² area for text-to-image, 1536² for editing. The sliders reach the longest
99
+ # side any preset has.
100
+ CUSTOM_CAP = {False: 2048, True: 1536} # keyed by "has references"
101
+ MAX_SIDE = max(max(size) for size in SIZES.values() if size)
102
 
103
  if STUDENT == "full":
104
  transformer = QwenImage21Transformer2DModel.from_pretrained(
 
174
 
175
 
176
  @gpu
177
+ def generate(prompt, image_1, image_2, image_3, size_label, seed, randomize_seed, enhance=True, steps=STEPS,
178
+ custom_width=1024, custom_height=1024):
179
  steps = min(max(int(steps), 3), 8) # 6 (RAW_NODES) is the validated schedule; 3-8 are accepted, fewer or more degrade quickly
180
  images = [image for image in (image_1, image_2, image_3) if image is not None]
181
  # unchecked "Randomize seed" always honours the box, so the UI never says one thing and does another
182
  seed = random.randint(0, MAX_SEED) if (randomize_seed or seed is None) else max(int(seed), 0)
183
  size = SIZES.get(size_label)
184
  width, height = size if size else (None, None)
185
+ ratio_name = size_label.split(" · ")[0]
186
+ clamped = ""
187
+ if size_label == CUSTOM:
188
+ cap = CUSTOM_CAP[bool(images)]
189
+ scale = min(1.0, cap / (custom_width * custom_height) ** 0.5)
190
+ width, height = int(custom_width * scale) // 32 * 32, int(custom_height * scale) // 32 * 32
191
+ ratio_name = f"{width}:{height}"
192
+ if scale < 1:
193
+ clamped = f" · scaled to {width}×{height} ({'editing' if images else 'text-to-image'} caps at {cap}² area)"
194
  # Belt and braces: an editing call must land on one of the editing buckets even if the caller
195
  # reaches generate() with a text-to-image size (stale dropdown, gr.Examples, API client). Test
196
  # membership in the offered dims rather than a raw area threshold: calculate_dimensions(1536**2,
197
  # 4/3) is 1760x1344 = 2_365_440 px, slightly *above* 1536**2, so an area test would flag two
198
  # sizes the editing menu itself offers. Snap to the 1536² bucket of the nearest named ratio.
199
+ elif images and width and (width, height) not in EDIT_DIMS:
 
200
  ratio = min(RATIOS, key=lambda name_ratio: abs(name_ratio[1] - width / height))[1]
201
  width, height, _ = calculate_dimensions(1536 * 1536, ratio)
202
  clamped = f" · clamped to {width}×{height} (editing caps at the 1536² bucket)"
203
  used_prompt, enhance_note = prompt, ""
204
  if enhance and prompt.strip():
205
  started = time.perf_counter()
206
+ used_prompt = enhance_prompt(prompt, images, ratio_name)
207
  enhance_note = f" · enhance {time.perf_counter() - started:.1f}s"
208
  if used_prompt == prompt:
209
  enhance_note += " (rewrite failed to parse, original prompt used)"
 
232
  return gr.update(choices=choices, value=current if current in choices else choices[0])
233
 
234
 
235
+ def custom_size_controls(size_label, width, height):
236
+ """The width/height sliders show only for Custom; a preset pre-fills them, so Custom starts from the last size picked."""
237
+ size = SIZES.get(size_label)
238
+ return gr.update(visible=size_label == CUSTOM), size[0] if size else width, size[1] if size else height
239
+
240
+
241
+ # Viggle brand tokens (viggle-brandbook): jolt-green calls to action with ink text, warm neutrals, Satoshi (from Fontshare)
242
+ JOLT = gr.themes.Color(c50="#e6fdec", c100="#c2f9d0", c200="#8ff2a9", c300="#52ea7c", c400="#1fe456", c500="#00e13f",
243
+ c600="#00b833", c700="#008f28", c800="#006b1e", c900="#004d16", c950="#002e0d", name="jolt")
244
+ WARM = gr.themes.Color(c50="#fcfcfc", c100="#f1eeea", c200="#e1e1e1", c300="#d9d9d9", c400="#9a9a9a", c500="#74685a",
245
+ c600="#4d453d", c700="#2e2a25", c800="#221d18", c900="#1b1614", c950="#151110", name="warm")
246
+ THEME = gr.themes.Default(primary_hue=JOLT, secondary_hue=JOLT, neutral_hue=WARM,
247
+ font=[gr.themes.Font("Satoshi"), "ui-sans-serif", "system-ui", "sans-serif"]).set(
248
+ body_text_color="#29231e", body_text_color_dark="#f4f1ec",
249
+ button_primary_background_fill="#00e13f", button_primary_background_fill_dark="#00e13f",
250
+ button_primary_background_fill_hover="#00c937", button_primary_background_fill_hover_dark="#00c937",
251
+ button_primary_border_color="#00e13f", button_primary_border_color_dark="#00e13f",
252
+ button_primary_text_color="#29231e", button_primary_text_color_dark="#29231e",
253
+ )
254
+ FONT_HEAD = '<link rel="stylesheet" href="https://api.fontshare.com/v2/css?f[]=satoshi@400,500,700&display=swap">'
255
+
256
+ with gr.Blocks(title="Viggle Turbo v0.2.1 · Qwen-Image-2.1 6-step", theme=THEME, head=FONT_HEAD) as demo:
257
  gr.Markdown(
258
  "# Viggle Turbo v0.2.1 — 6-step Qwen-Image-2.1\n"
259
+ "Text-to-image and image editing in **6 steps**: about **5× faster** than the 40-step Qwen-Image-2.1 and very "
260
+ "competitive with it in quality — see the **Comparison** tab. Leave the references empty for text-to-image, "
261
+ "or add one to three to edit, compose or transfer style."
 
 
 
 
 
 
 
262
  )
263
+ with gr.Accordion("About this model", open=False):
264
+ gr.Markdown(
265
+ "A DMD-distilled student of **Qwen-Image-2.1**, sampled with no classifier-free guidance. On some prompts we "
266
+ "prefer its output to the base model's; complicated edits (multi-reference composition, face swaps, "
267
+ "identity-preserving edits) can still fall short of it. The **Comparison** tab has 33 examples of the official "
268
+ "Qwen Space side by side, turbo vs base, same seed.\n\n"
269
+ "**v0.2.1 (2026-09-24):** the step-700 checkpoint of the v0.2 run on the 6-step schedule — intra-prompt diversity "
270
+ "0.98× the base model, 0% composition drift; earlier versions and the numbers are on the model card. "
271
+ f"Weights: **{STUDENT_TAG}** from [{STUDENT_REPO}](https://huggingface.co/{STUDENT_REPO})."
272
+ )
273
  with gr.Tab("Generate"):
274
  with gr.Row():
275
  with gr.Column(scale=3):
 
278
  image_1 = gr.Image(label="Reference 1 (optional)", type="pil", image_mode="RGB", height=200)
279
  image_2 = gr.Image(label="Reference 2 (optional)", type="pil", image_mode="RGB", height=200)
280
  image_3 = gr.Image(label="Reference 3 (optional)", type="pil", image_mode="RGB", height=200)
281
+ # allow_custom_value: gradio validates API calls against the *initial* choices (the text-to-image
282
+ # menu), which would reject every editing size sent through /generate; generate() snaps
283
+ # anything off-menu itself.
284
+ size_label = gr.Dropdown(label="Output size", choices=T2I_CHOICES, value=T2I_CHOICES[0], allow_custom_value=True,
285
+ info="The menu switches to the editing sizes as soon as a reference is attached. Custom "
286
+ "sizes above 2048² area (1536² when editing) are scaled down, keeping the ratio.")
287
+ with gr.Row(visible=False) as custom_row:
288
+ custom_width = gr.Slider(label="Width", minimum=512, maximum=MAX_SIDE, step=32, value=1024)
289
+ custom_height = gr.Slider(label="Height", minimum=512, maximum=MAX_SIDE, step=32, value=1024)
 
290
  enhance = gr.Checkbox(
291
  label="Enhance prompt",
292
  value=True,
293
  info="Rewrite the prompt with the official Qwen-Image prompt-enhancement instructions "
294
  "(runs on the built-in Qwen3-VL text encoder, +4–15 s). The text actually sent is shown under the result.",
295
  )
296
+ with gr.Accordion("Advanced settings", open=False) as advanced:
297
+ steps = gr.Slider(label="Steps", minimum=3, maximum=8, step=1, value=STEPS,
298
+ info="6 is the validated schedule; extra steps are added at the high-noise end.")
299
+ with gr.Row():
300
+ seed = gr.Number(label="Seed", value=0, precision=0, interactive=False)
301
+ randomize_seed = gr.Checkbox(label="Randomize seed", value=True)
302
+ run = gr.Button("Generate", variant="primary", size="lg")
303
  with gr.Column(scale=4):
304
  output_image = gr.Image(label="Result", type="pil", height=560)
305
  info = gr.Markdown()
 
307
 
308
  def example_details(prompt, image_1, image_2, image_3, result):
309
  """Runs on an example click (no GPU): fills the status line and the prompt stored with the row, and sets the
310
+ steps / size / enhancement / seed controls to what produced it, so Generate reproduces the result. It opens
311
+ Advanced settings to show that the seed is now fixed. It also sets the size menu's mode itself: loading an
312
+ example does not fire the reference images' .input listeners."""
313
  row = next(row for row in EXAMPLES if row["prompt"] == prompt)
314
  choices = EDIT_CHOICES if any(row["refs"]) else T2I_CHOICES
315
+ return (row["info"], row["used_prompt"], row["steps"], gr.update(choices=choices, value=row["size_label"]),
316
+ row["enhance"], row["seed"], False, gr.update(open=True))
317
 
318
  if EXAMPLES:
319
  gr.Examples(
 
323
  ],
324
  inputs=[prompt, image_1, image_2, image_3, output_image],
325
  fn=example_details,
326
+ outputs=[info, used_prompt, steps, size_label, enhance, seed, randomize_seed, advanced],
327
  run_on_click=True,
328
+ # On Spaces gr.Examples caches by default (lazily on ZeroGPU), and gradio 5.50 crashes caching an output that
329
+ # updates a Dropdown's choices ('Dropdown' object has no attribute 'proxy_url'). example_details needs no GPU.
330
+ cache_examples=False,
331
  examples_per_page=12,
332
+ label="Examples · results pre-rendered by this model — click a row to load it with its settings; Generate reproduces it",
333
  )
334
 
335
  with gr.Tab("Comparison"):
 
352
  inputs=[image_1, image_2, image_3, size_label],
353
  outputs=size_label,
354
  )
355
+ size_label.change(custom_size_controls, inputs=[size_label, custom_width, custom_height],
356
+ outputs=[custom_row, custom_width, custom_height])
357
  randomize_seed.change(lambda on: gr.update(interactive=not on), inputs=randomize_seed, outputs=seed)
358
  run.click(
359
  generate,
360
+ inputs=[prompt, image_1, image_2, image_3, size_label, seed, randomize_seed, enhance, steps, custom_width, custom_height],
361
  outputs=[output_image, info, used_prompt],
362
  )
363