yycc commited on
Commit
6ccfa08
·
verified ·
1 Parent(s): 77e9b3a

v0.2.1: default LoRA -> v6_isg step-700 EMA r256, examples re-rendered

Browse files
NOTICE CHANGED
@@ -2,7 +2,7 @@ Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 H
2
 
3
  Built with Qwen.
4
 
5
- This Space demonstrates Viggle/Qwen-Image-2.1-viggle-turbo, a DMD-distilled 4-step LoRA
6
  derivative of Qwen/Qwen-Image-2.1. The base weights are downloaded at runtime from
7
  Qwen/Qwen-Image-2.1 and are unmodified. Use is limited to non-commercial (research or
8
  evaluation) purposes, per the Qwen RESEARCH LICENSE AGREEMENT included as LICENSE.
 
2
 
3
  Built with Qwen.
4
 
5
+ This Space demonstrates Viggle/Qwen-Image-2.1-viggle-turbo, a DMD-distilled 6-step LoRA
6
  derivative of Qwen/Qwen-Image-2.1. The base weights are downloaded at runtime from
7
  Qwen/Qwen-Image-2.1 and are unmodified. Use is limited to non-commercial (research or
8
  evaluation) purposes, per the Qwen RESEARCH LICENSE AGREEMENT included as LICENSE.
README.md CHANGED
@@ -1,5 +1,5 @@
1
  ---
2
- title: Viggle Turbo v0.2 - 6-step Qwen-Image-2.1
3
  emoji: ⚡
4
  colorFrom: indigo
5
  colorTo: purple
@@ -11,7 +11,7 @@ pinned: false
11
  license: other
12
  license_name: qwen-research
13
  license_link: ./LICENSE
14
- short_description: v0.2 preview - 6-step Qwen-Image-2.1, T2I + editing
15
  models:
16
  - Viggle/Qwen-Image-2.1-viggle-turbo
17
  - Qwen/Qwen-Image-2.1
@@ -28,7 +28,7 @@ tags:
28
  # `hf_oauth` is not needed: the app calls no user-scoped Hub API.
29
  ---
30
 
31
- # Viggle Turbo v0.2 (preview) — 6-step Qwen-Image-2.1
32
 
33
  A DMD-distilled student of **Qwen-Image-2.1** that generates and edits images in **6 sampling steps**
34
  with **no classifier-free guidance**, against the teacher's 40 steps.
@@ -40,7 +40,11 @@ for the numbers).
40
  **2026-09-24: 6 steps instead of 5, same weights.** The extra step splits the highest-noise segment once more
41
  (raw nodes `1, 0.9375, 0.875` instead of `1, 0.875`; the low-noise nodes are unchanged). Composition drift
42
  against the base model goes from 4% of prompts to 0%, diversity from 0.93× to 0.97×, and the images are a
43
- little sharper. The examples below are re-rendered on the 6-step schedule.
 
 
 
 
44
 
45
  One model does both tasks, exactly as the base does: leave the reference slots empty for
46
  text-to-image, or fill one to three of them for editing / composition / style transfer.
@@ -57,16 +61,16 @@ three-reference prompt come from the
57
 
58
  **Built with Qwen.** Distilled from [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1).
59
 
60
- > **Status: preview, work in progress.** v0.2 is a large step up from v0.1 but still falls short of the 40-step base
61
  > model on complicated image editing (multi-reference composition, face swaps, identity-preserving edits, instructions
62
  > with several constraints). We will keep updating this Space as the distillation improves.
63
 
64
  ## What the app loads
65
 
66
- The Space runs the **v0.2 LoRA (rank 256)** by default, loaded at runtime on top of the base transformer with
67
  `pipe.load_lora_weights(STUDENT_REPO, weight_name=LORA_FILE)` and never merged (merging into bf16 keeps only ~47 %
68
- of the adapter delta). `LORA_FILE` defaults to `Qwen-Image-2.1-viggle-turbo-v0.2-5step-lora-r256.safetensors`; the
69
- v0.1 r64 adapter is still in the model repo and can be selected with it. The v0.1 **full fine-tuned transformer**
70
  is still selectable with `STUDENT=full`:
71
 
72
  ```python
@@ -186,7 +190,7 @@ rewrite adds 2-8 s on a B200; `duration=90` covers both.
186
  **Cold start** is dominated by the download: ~30.9 GiB of base safetensors from
187
  `Qwen/Qwen-Image-2.1` (13.25 transformer + 16.33 text encoder + 1.26 VAE + 0.02 processor) plus
188
  **1.3 GiB** for the LoRA — the shipped root adapter is bf16, and `load_lora_weights(weight_name=…)`
189
- fetches that single file, not the F32 `peft_v0.2/` copy — then ~30–40 s to load and pack the pipeline
190
  onto the GPU. Only the adapter changes
191
  between releases, so a cached Space restarts far faster than it first boots.
192
 
 
1
  ---
2
+ title: Viggle Turbo v0.2.1 - 6-step Qwen-Image-2.1
3
  emoji: ⚡
4
  colorFrom: indigo
5
  colorTo: purple
 
11
  license: other
12
  license_name: qwen-research
13
  license_link: ./LICENSE
14
+ short_description: v0.2.1 preview - 6-step Qwen-Image-2.1, T2I + editing
15
  models:
16
  - Viggle/Qwen-Image-2.1-viggle-turbo
17
  - Qwen/Qwen-Image-2.1
 
28
  # `hf_oauth` is not needed: the app calls no user-scoped Hub API.
29
  ---
30
 
31
+ # Viggle Turbo v0.2.1 (preview) — 6-step Qwen-Image-2.1
32
 
33
  A DMD-distilled student of **Qwen-Image-2.1** that generates and edits images in **6 sampling steps**
34
  with **no classifier-free guidance**, against the teacher's 40 steps.
 
40
  **2026-09-24: 6 steps instead of 5, same weights.** The extra step splits the highest-noise segment once more
41
  (raw nodes `1, 0.9375, 0.875` instead of `1, 0.875`; the low-noise nodes are unchanged). Composition drift
42
  against the base model goes from 4% of prompts to 0%, diversity from 0.93× to 0.97×, and the images are a
43
+ little sharper.
44
+
45
+ **v0.2.1 (2026-09-24):** the step-700 checkpoint of the same training run (v0.2 was step 600): a little
46
+ sharper and marginally more diverse (0.98× the base model), still 0% composition drift. The examples below are
47
+ rendered with it on the 6-step schedule.
48
 
49
  One model does both tasks, exactly as the base does: leave the reference slots empty for
50
  text-to-image, or fill one to three of them for editing / composition / style transfer.
 
61
 
62
  **Built with Qwen.** Distilled from [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1).
63
 
64
+ > **Status: preview, work in progress.** v0.2.1 is a large step up from v0.1 but still falls short of the 40-step base
65
  > model on complicated image editing (multi-reference composition, face swaps, identity-preserving edits, instructions
66
  > with several constraints). We will keep updating this Space as the distillation improves.
67
 
68
  ## What the app loads
69
 
70
+ The Space runs the **v0.2.1 LoRA (rank 256)** by default, loaded at runtime on top of the base transformer with
71
  `pipe.load_lora_weights(STUDENT_REPO, weight_name=LORA_FILE)` and never merged (merging into bf16 keeps only ~47 %
72
+ of the adapter delta). `LORA_FILE` defaults to `Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r256.safetensors`; the
73
+ v0.2 and v0.1 adapters are still in the model repo and can be selected with it. The v0.1 **full fine-tuned transformer**
74
  is still selectable with `STUDENT=full`:
75
 
76
  ```python
 
190
  **Cold start** is dominated by the download: ~30.9 GiB of base safetensors from
191
  `Qwen/Qwen-Image-2.1` (13.25 transformer + 16.33 text encoder + 1.26 VAE + 0.02 processor) plus
192
  **1.3 GiB** for the LoRA — the shipped root adapter is bf16, and `load_lora_weights(weight_name=…)`
193
+ fetches that single file, not the F32 `peft_v0.2.1/` copy — then ~30–40 s to load and pack the pipeline
194
  onto the GPU. Only the adapter changes
195
  between releases, so a cached Space restarts far faster than it first boots.
196
 
app.py CHANGED
@@ -1,4 +1,4 @@
1
- """Gradio demo for the DMD-distilled Qwen-Image-2.1 student (v0.2: LoRA r256 sampled in 6 steps; v0.1 students still selectable)."""
2
 
3
  # ZeroGPU patches torch at import time, so `spaces` must be imported before torch or any CUDA use.
4
  # find_spec keeps this file runnable off-Spaces, where the package is absent.
@@ -32,18 +32,18 @@ from PIL import Image
32
  from diffusers import FlowMatchEulerDiscreteScheduler, QwenImage21Pipeline, QwenImage21Transformer2DModel
33
  from diffusers.pipelines.qwenimage21.pipeline_qwenimage21 import calculate_dimensions
34
 
35
- # Students in the model repo. STUDENT="lora" (default): the v0.2 LoRA (r256) at the repo root, applied at runtime on
36
  # top of the base transformer and never merged (merging into bf16 keeps only ~47% of the adapter delta, the
37
- # per-element deltas sit below the bf16 ULP of the base weights). LORA_FILE selects the adapter file; the v0.1
38
- # r64 adapter is still in the repo. STUDENT="full": the v0.1 full fine-tune in transformer/ (bf16, exact, no
39
  # adapter) replaces the base transformer at load time - the base transformer is then not downloaded at all.
40
  BASE_MODEL_ID = os.environ.get("BASE_MODEL_ID", "Qwen/Qwen-Image-2.1")
41
  STUDENT_REPO = os.environ.get("STUDENT_REPO", os.environ.get("LORA_REPO", "Viggle/Qwen-Image-2.1-viggle-turbo"))
42
  STUDENT = os.environ.get("STUDENT", "lora")
43
- LORA_FILE = os.environ.get("LORA_FILE", "Qwen-Image-2.1-viggle-turbo-v0.2-5step-lora-r256.safetensors")
44
- STUDENT_TAG = {"full": "v0.1 full fine-tune (transformer/)", "lora": "v0.2 LoRA r256" if "v0.2" in LORA_FILE else f"LoRA {LORA_FILE}"}[STUDENT]
45
  HF_TOKEN = os.environ.get("HF_TOKEN") # only needed while the model repo is private
46
- # v0.2 is sampled on 6 Euler steps whose raw (pre-shift) sigma nodes are RAW_NODES: the 4-step training schedule
47
  # linspace(1, 1/4, 4) with its highest-noise segment [1, 0.75] cut into three (1, 0.9375, 0.875, 0.75). The composition
48
  # is decided in that segment, and one big Euler step there ghosts and drifts the layout; the low-noise nodes 0.75, 0.5,
49
  # 0.25 are the ones the student was trained to land on, and moving them softens the image. So every extra step goes to the high-noise
@@ -216,9 +216,9 @@ def refresh_sizes(image_1, image_2, image_3, current):
216
  return gr.update(choices=choices, value=current if current in choices else choices[0])
217
 
218
 
219
- with gr.Blocks(title="Viggle Turbo v0.2 · Qwen-Image-2.1 6-step") as demo:
220
  gr.Markdown(
221
- "# Viggle Turbo v0.2 (preview) — 6-step Qwen-Image-2.1\n"
222
  "A DMD-distilled student of **Qwen-Image-2.1** that generates and edits in **6 sampling steps**, "
223
  "with no classifier-free guidance. Leave the reference images empty for text-to-image; "
224
  "add one to three of them to edit, compose or transfer style. The size menu switches to the "
@@ -226,7 +226,8 @@ with gr.Blocks(title="Viggle Turbo v0.2 · Qwen-Image-2.1 6-step") as demo:
226
  "**v0.2:** much better sample diversity than v0.1 (intra-prompt diversity 0.97× the 40-step base model, up from 0.75×) "
227
  "and closer prompt adherence / composition to the base model. **2026-09-24:** same weights, now sampled on 6 steps "
228
  "instead of 5 - the extra step splits the highest-noise segment again, which removes the last of the composition drift "
229
- "(0% of prompts vs 4% at 5 steps) and lifts diversity from 0.93× to 0.97×. Still a preview: complicated edits "
 
230
  "(multi-reference composition, face swaps, identity-preserving edits) remain weaker than the 40-step base model.\n\n"
231
  f"Weights: **{STUDENT_TAG}** from "
232
  f"[{STUDENT_REPO}](https://huggingface.co/{STUDENT_REPO})."
 
1
+ """Gradio demo for the DMD-distilled Qwen-Image-2.1 student (v0.2.1: LoRA r256 sampled in 6 steps; v0.2 / v0.1 students still selectable)."""
2
 
3
  # ZeroGPU patches torch at import time, so `spaces` must be imported before torch or any CUDA use.
4
  # find_spec keeps this file runnable off-Spaces, where the package is absent.
 
32
  from diffusers import FlowMatchEulerDiscreteScheduler, QwenImage21Pipeline, QwenImage21Transformer2DModel
33
  from diffusers.pipelines.qwenimage21.pipeline_qwenimage21 import calculate_dimensions
34
 
35
+ # Students in the model repo. STUDENT="lora" (default): the v0.2.1 LoRA (r256) at the repo root, applied at runtime on
36
  # top of the base transformer and never merged (merging into bf16 keeps only ~47% of the adapter delta, the
37
+ # per-element deltas sit below the bf16 ULP of the base weights). LORA_FILE selects the adapter file; the v0.2 and
38
+ # v0.1 adapters are still in the repo. STUDENT="full": the v0.1 full fine-tune in transformer/ (bf16, exact, no
39
  # adapter) replaces the base transformer at load time - the base transformer is then not downloaded at all.
40
  BASE_MODEL_ID = os.environ.get("BASE_MODEL_ID", "Qwen/Qwen-Image-2.1")
41
  STUDENT_REPO = os.environ.get("STUDENT_REPO", os.environ.get("LORA_REPO", "Viggle/Qwen-Image-2.1-viggle-turbo"))
42
  STUDENT = os.environ.get("STUDENT", "lora")
43
+ LORA_FILE = os.environ.get("LORA_FILE", "Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r256.safetensors")
44
+ STUDENT_TAG = {"full": "v0.1 full fine-tune (transformer/)", "lora": "v0.2.1 LoRA r256" if "v0.2.1" in LORA_FILE else f"LoRA {LORA_FILE}"}[STUDENT]
45
  HF_TOKEN = os.environ.get("HF_TOKEN") # only needed while the model repo is private
46
+ # The LoRA is sampled on 6 Euler steps whose raw (pre-shift) sigma nodes are RAW_NODES: the 4-step training schedule
47
  # linspace(1, 1/4, 4) with its highest-noise segment [1, 0.75] cut into three (1, 0.9375, 0.875, 0.75). The composition
48
  # is decided in that segment, and one big Euler step there ghosts and drifts the layout; the low-noise nodes 0.75, 0.5,
49
  # 0.25 are the ones the student was trained to land on, and moving them softens the image. So every extra step goes to the high-noise
 
216
  return gr.update(choices=choices, value=current if current in choices else choices[0])
217
 
218
 
219
+ with gr.Blocks(title="Viggle Turbo v0.2.1 · Qwen-Image-2.1 6-step") as demo:
220
  gr.Markdown(
221
+ "# Viggle Turbo v0.2.1 (preview) — 6-step Qwen-Image-2.1\n"
222
  "A DMD-distilled student of **Qwen-Image-2.1** that generates and edits in **6 sampling steps**, "
223
  "with no classifier-free guidance. Leave the reference images empty for text-to-image; "
224
  "add one to three of them to edit, compose or transfer style. The size menu switches to the "
 
226
  "**v0.2:** much better sample diversity than v0.1 (intra-prompt diversity 0.97× the 40-step base model, up from 0.75×) "
227
  "and closer prompt adherence / composition to the base model. **2026-09-24:** same weights, now sampled on 6 steps "
228
  "instead of 5 - the extra step splits the highest-noise segment again, which removes the last of the composition drift "
229
+ "(0% of prompts vs 4% at 5 steps) and lifts diversity from 0.93× to 0.97×. **v0.2.1 (2026-09-24):** the step-700 "
230
+ "checkpoint of the same run (v0.2 was step 600) - a little sharper, diversity 0.98×, still 0% drift. Still a preview: complicated edits "
231
  "(multi-reference composition, face swaps, identity-preserving edits) remain weaker than the 40-step base model.\n\n"
232
  f"Weights: **{STUDENT_TAG}** from "
233
  f"[{STUDENT_REPO}](https://huggingface.co/{STUDENT_REPO})."
examples/edit_2ref_cat.png CHANGED

Git LFS Details

  • SHA256: 7bdb9799c0a46be71d9267e43c7ed7d468f685bc52b5b2e6c5046395e1f1fe23
  • Pointer size: 132 Bytes
  • Size of remote file: 2.08 MB

Git LFS Details

  • SHA256: 9ae7cad5b1214eddaa1da535b77b39e0f467be731a52b0d57d0f3c3a0f0b1461
  • Pointer size: 132 Bytes
  • Size of remote file: 2.08 MB
examples/edit_3ref_klein.png CHANGED

Git LFS Details

  • SHA256: af2b19c191693c0389ac10a6d24f02381bdbb12a3d49bb9f491f7d64562d1b25
  • Pointer size: 132 Bytes
  • Size of remote file: 1.69 MB

Git LFS Details

  • SHA256: 69eed69031bad070ec2d03c97fb8cc8ff8668f262c16ac492fa734d796f0d8be
  • Pointer size: 132 Bytes
  • Size of remote file: 1.69 MB
examples/edit_sketch.png CHANGED

Git LFS Details

  • SHA256: ba5c1d9b1fcdb8c45866c7848f2ba3fba2390849ea622013e20b34b8df2276f2
  • Pointer size: 132 Bytes
  • Size of remote file: 2.46 MB

Git LFS Details

  • SHA256: ab32269a4f96d2f6553766b0740a3b2aebb1d61ca535ba0f568b6802262c6583
  • Pointer size: 132 Bytes
  • Size of remote file: 2.48 MB
examples/manifest.json CHANGED
@@ -7,7 +7,7 @@
7
  "bird.webp"
8
  ],
9
  "result": "edit_3ref_klein.png",
10
- "info": "Pre-rendered example · seed `42` · 928×1152 · 6 steps · v0.2 LoRA r256 · prompt enhanced",
11
  "used_prompt": "Compose a scene where the person from <image1>, standing beside the vintage brown car, is gently petting the fluffy cat from <image2> that is perched on a stone windowsill with green shutters, while the bird from <image3>, with a red crown and long beak, stands next to them, all under warm cinematic lighting with a shallow depth of field focusing on the trio, keeping the background and subjects' identities from the input images unchanged.",
12
  "steps": 6,
13
  "size_label": "Auto · match the last reference (1024² area)",
@@ -21,7 +21,7 @@
21
  null
22
  ],
23
  "result": "t2i_launch.png",
24
- "info": "Pre-rendered example · seed `3` · 2496×1664 · 8 steps · v0.2 LoRA r256 · prompt sent as written (enhancement off)",
25
  "used_prompt": "Create a premium, minimalist launch graphic for an AI image model, designed for a Twitter/X announcement.\n\nAspect ratio: 16:9.\nStyle: elegant, high-end, modern AI product launch visual. Clean Apple / NVIDIA / premium creative-software aesthetic. Dark, cinematic, sophisticated, not flashy or cluttered.\n\nBACKGROUND\n- Deep black to charcoal gradient background.\n- Very subtle blue and violet ambient glow.\n- Soft cinematic lighting, slight glossy reflections near the bottom.\n- Large areas of negative space.\n- No unnecessary particles, grids, icons, decorative UI, or busy textures.\n\nTOP-LEFT TYPOGRAPHY\n\nMain title:\n“Qwen-Image-2.1 ❤️ Viggle Turbo”\n\n- Large bold geometric sans-serif.\n- Clean Swiss-style typography.\n- Bright soft-white text.\n- Red heart between the two model names.\n- Precise kerning and professional spacing.\n- Keep the title on one line.\n- Position around 4–5% from the left and 8–10% from the top.\n\nDirectly underneath, smaller subtitle:\n“DMD-distilled · Generate & Edit”\n\n- Thin or regular sans-serif.\n- Around 30–35% of the main title font size.\n- Soft light-gray / off-white.\n- Generous letter spacing.\n- Minimal and understated.\n\nCENTER HERO MESSAGE\n\nLarge dominant typography:\n\n“6-step”\n“generation”\n\non two lines.\n\n- “6-step” should be extremely large.\n- “generation” slightly smaller but still bold.\n- Heavy geometric sans-serif.\n- White / very subtle cool-white gradient.\n- Tight line spacing.\n- Perfectly clean typography.\n- Position slightly above the vertical center.\n- The text should be the strongest visual element in the design.\n- No exaggerated 3D text, extrusion, chrome effects, or heavy shadows.\n- Only a very subtle soft glow.\n\nVISUAL ELEMENT\n\nCreate a single elegant curved ribbon of generated images flowing across the lower half of the composition.\n\nThe ribbon should:\n- Start from the lower-left foreground.\n- Curve smoothly toward the center-right.\n- Continue upward toward the upper-right background.\n- Feel like a sophisticated cinematic film strip or flowing image-generation sequence.\n- Use approximately 5–7 image panels only.\n- Avoid a dense collage.\n\nEach panel:\n- Rounded rectangle.\n- Thin subtle border.\n- Premium glossy display appearance.\n- Slight perspective distortion following the curve.\n- Gradually decrease in size toward the background.\n- Use realistic generated landscape photography:\n • alpine lake\n • mountain valley\n • waterfall\n • dramatic mountain peaks\n • coastal sunset\n- Rich but natural colors.\n- Foreground panels sharp, distant panels progressively softer / slightly blurred.\n- Strong depth-of-field.\n\nThe ribbon itself should have an extremely subtle luminous edge:\n- cyan / electric blue on the left\n- transitioning gently toward violet / magenta on the right\n\nKeep the glow restrained and premium.\nDo not make it look like cyberpunk neon.\n\nLIGHTING\n\nAdd one very thin horizontal blue-to-violet light flare passing subtly behind the “6-step” text.\n\nUse soft reflected blue/violet light underneath the image ribbon.\n\nLighting should feel cinematic and expensive, not game-like.\n\nBOTTOM-RIGHT\n\nSmall footer:\n“Generate with viggle-turbo”\n\n- Small clean sans-serif.\n- White / light gray.\n- Bottom-right aligned.\n- Approximately 5% from the right and 6–8% from the bottom.\n- Lots of breathing room.\n- Do not make it prominent.\n\nOVERALL COMPOSITION\n\nThe visual hierarchy should clearly be:\n\n1. “6-step generation”\n2. “Qwen-Image-2.1 ❤️ Viggle Turbo”\n3. flowing generated-image ribbon\n4. “DMD-distilled · Generate & Edit”\n5. small “Generate with viggle-turbo.” footer\n\nThe graphic should immediately communicate:\nfast generation, image generation, premium AI model, six-step inference.\n\nOverall aesthetic:\nminimal, elegant, expensive, technically sophisticated, cinematic, polished, restrained, premium AI launch campaign.\n\nAvoid:\n- excessive neon\n- cyberpunk look\n- too many thumbnails\n- busy collage layouts\n- random icons\n- fake UI elements\n- gradients inside every object\n- excessive lens flares\n- heavy shadows\n- 3D metallic typography\n- cartoon imagery\n- excessive text\n- logos other than the specified text\n- watermark\n- spelling errors\n- distorted letters\n- duplicated text",
26
  "steps": 8,
27
  "size_label": "3:2 · 2496×1664 (2048² area)",
@@ -35,7 +35,7 @@
35
  null
36
  ],
37
  "result": "t2i_capybara.png",
38
- "info": "Pre-rendered example · seed `0` · 1248×832 · 6 steps · v0.2 LoRA r256 · prompt enhanced",
39
  "used_prompt": "The image is a close-up realistic photograph of a soaking wet capybara taking shelter under a large banana leaf in a rainy jungle, the background a lush green canopy dappled with falling raindrops and misty humidity. In the upper-left corner, the edge of the banana leaf curves downward, its broad surface glistening with rainwater and casting a soft shadow over the capybara’s back. Across the top third, rain streaks blur the dense foliage behind, with shafts of diffused light filtering through the canopy. In the center of the frame, the capybara’s rounded body is pressed against the leaf’s underside, its fur matted and dark with water, clinging to its skin in clumps; its small ears are flattened against its head, and its eyes are half-closed, glistening with moisture. On the right side, a second banana leaf peeks into the frame, its veins visible and slightly curled at the edges, partially obscuring the background. Behind the capybara’s left shoulder, a cluster of broad, wet leaves hangs low, their edges blurred by the shallow depth of field. Along the lower edge, the ground is a muddy brown, slick with rainwater and scattered with fallen leaves and twigs, reflecting the dim light. The lighting is soft, diffused daylight filtered through the jungle canopy, casting gentle highlights on the capybara’s wet fur and the water droplets clinging to the banana leaf, with faint shadows beneath the animal and along the leaf’s underside. The overall composition feels intimate and sheltered, emphasizing the capybara’s vulnerability and the quiet resilience of nature in the downpour.",
40
  "steps": 6,
41
  "size_label": "3:2 · 1248×832 (1024² area)",
@@ -49,7 +49,7 @@
49
  null
50
  ],
51
  "result": "t2i_diorama.png",
52
- "info": "Pre-rendered example · seed `0` · 2368×1760 · 6 steps · v0.2 LoRA r256 · prompt enhanced",
53
  "used_prompt": "The image is a 4:3 isometric miniature diorama of a mountain research station, rendered in high detail with soft studio lighting and a tilt-shift effect. The scene is set on a rugged, snow-dusted mountain slope, with the research station occupying the central foreground, its modular buildings and domed observatory nestled into the terrain. To the left, a steep incline rises into a misty valley, while the right side features a rocky outcrop with a small antenna array. Behind the station, a jagged mountain ridge stretches across the background under a pale blue sky with soft, diffused clouds. The station’s structures are detailed with weathered metal siding, glass windows reflecting the ambient light, and exposed piping and cables. A small helipad with a landing pad marker is visible in the lower-left corner, and a narrow service road winds up the slope toward the main entrance. The lighting is soft and even, emanating from an unseen source above and slightly to the left, casting gentle shadows that emphasize the three-dimensional depth of the diorama. The overall composition feels immersive and meticulously crafted, with a balanced arrangement of architectural elements and natural terrain, evoking a sense of isolation and scientific exploration in a remote alpine environment.",
54
  "steps": 6,
55
  "size_label": "4:3 · 2368×1760 (2048² area)",
@@ -63,7 +63,7 @@
63
  null
64
  ],
65
  "result": "edit_sketch.png",
66
- "info": "Pre-rendered example · seed `3` · 1024×1024 · 6 steps · v0.2 LoRA r256 · prompt enhanced",
67
  "used_prompt": "Convert the image to a pencil sketch on textured paper, preserving the exact composition, framing, and spatial arrangement of all elements including the rain-streaked window, the interior bookstore shelves, the figure inside, and the \"OPEN LATE\" sign in the foreground, while rendering the entire scene in a hand-drawn, sketchy style with visible pencil strokes and paper texture.",
68
  "steps": 6,
69
  "size_label": "Auto · match the last reference (1024² area)",
@@ -77,7 +77,7 @@
77
  null
78
  ],
79
  "result": "edit_2ref_cat.png",
80
- "info": "Pre-rendered example · seed `42` · 832×1248 · 6 steps · v0.2 LoRA r256 · prompt enhanced",
81
  "used_prompt": "Compose a new image where the woman from <image1>, wearing her full traditional outfit with skull makeup and holding a smoke bomb, is now holding the fluffy cat from <image2> in her arms, positioned in the same spot as in the original image, with the ornate wooden gate and surrounding smoke from <image1> remaining unchanged as the background.",
82
  "steps": 6,
83
  "size_label": "Auto · match the last reference (1024² area)",
 
7
  "bird.webp"
8
  ],
9
  "result": "edit_3ref_klein.png",
10
+ "info": "Pre-rendered example · seed `42` · 928×1152 · 6 steps · v0.2.1 LoRA r256 · prompt enhanced",
11
  "used_prompt": "Compose a scene where the person from <image1>, standing beside the vintage brown car, is gently petting the fluffy cat from <image2> that is perched on a stone windowsill with green shutters, while the bird from <image3>, with a red crown and long beak, stands next to them, all under warm cinematic lighting with a shallow depth of field focusing on the trio, keeping the background and subjects' identities from the input images unchanged.",
12
  "steps": 6,
13
  "size_label": "Auto · match the last reference (1024² area)",
 
21
  null
22
  ],
23
  "result": "t2i_launch.png",
24
+ "info": "Pre-rendered example · seed `3` · 2496×1664 · 8 steps · v0.2.1 LoRA r256 · prompt sent as written (enhancement off)",
25
  "used_prompt": "Create a premium, minimalist launch graphic for an AI image model, designed for a Twitter/X announcement.\n\nAspect ratio: 16:9.\nStyle: elegant, high-end, modern AI product launch visual. Clean Apple / NVIDIA / premium creative-software aesthetic. Dark, cinematic, sophisticated, not flashy or cluttered.\n\nBACKGROUND\n- Deep black to charcoal gradient background.\n- Very subtle blue and violet ambient glow.\n- Soft cinematic lighting, slight glossy reflections near the bottom.\n- Large areas of negative space.\n- No unnecessary particles, grids, icons, decorative UI, or busy textures.\n\nTOP-LEFT TYPOGRAPHY\n\nMain title:\n“Qwen-Image-2.1 ❤️ Viggle Turbo”\n\n- Large bold geometric sans-serif.\n- Clean Swiss-style typography.\n- Bright soft-white text.\n- Red heart between the two model names.\n- Precise kerning and professional spacing.\n- Keep the title on one line.\n- Position around 4–5% from the left and 8–10% from the top.\n\nDirectly underneath, smaller subtitle:\n“DMD-distilled · Generate & Edit”\n\n- Thin or regular sans-serif.\n- Around 30–35% of the main title font size.\n- Soft light-gray / off-white.\n- Generous letter spacing.\n- Minimal and understated.\n\nCENTER HERO MESSAGE\n\nLarge dominant typography:\n\n“6-step”\n“generation”\n\non two lines.\n\n- “6-step” should be extremely large.\n- “generation” slightly smaller but still bold.\n- Heavy geometric sans-serif.\n- White / very subtle cool-white gradient.\n- Tight line spacing.\n- Perfectly clean typography.\n- Position slightly above the vertical center.\n- The text should be the strongest visual element in the design.\n- No exaggerated 3D text, extrusion, chrome effects, or heavy shadows.\n- Only a very subtle soft glow.\n\nVISUAL ELEMENT\n\nCreate a single elegant curved ribbon of generated images flowing across the lower half of the composition.\n\nThe ribbon should:\n- Start from the lower-left foreground.\n- Curve smoothly toward the center-right.\n- Continue upward toward the upper-right background.\n- Feel like a sophisticated cinematic film strip or flowing image-generation sequence.\n- Use approximately 5–7 image panels only.\n- Avoid a dense collage.\n\nEach panel:\n- Rounded rectangle.\n- Thin subtle border.\n- Premium glossy display appearance.\n- Slight perspective distortion following the curve.\n- Gradually decrease in size toward the background.\n- Use realistic generated landscape photography:\n • alpine lake\n • mountain valley\n • waterfall\n • dramatic mountain peaks\n • coastal sunset\n- Rich but natural colors.\n- Foreground panels sharp, distant panels progressively softer / slightly blurred.\n- Strong depth-of-field.\n\nThe ribbon itself should have an extremely subtle luminous edge:\n- cyan / electric blue on the left\n- transitioning gently toward violet / magenta on the right\n\nKeep the glow restrained and premium.\nDo not make it look like cyberpunk neon.\n\nLIGHTING\n\nAdd one very thin horizontal blue-to-violet light flare passing subtly behind the “6-step” text.\n\nUse soft reflected blue/violet light underneath the image ribbon.\n\nLighting should feel cinematic and expensive, not game-like.\n\nBOTTOM-RIGHT\n\nSmall footer:\n“Generate with viggle-turbo”\n\n- Small clean sans-serif.\n- White / light gray.\n- Bottom-right aligned.\n- Approximately 5% from the right and 6–8% from the bottom.\n- Lots of breathing room.\n- Do not make it prominent.\n\nOVERALL COMPOSITION\n\nThe visual hierarchy should clearly be:\n\n1. “6-step generation”\n2. “Qwen-Image-2.1 ❤️ Viggle Turbo”\n3. flowing generated-image ribbon\n4. “DMD-distilled · Generate & Edit”\n5. small “Generate with viggle-turbo.” footer\n\nThe graphic should immediately communicate:\nfast generation, image generation, premium AI model, six-step inference.\n\nOverall aesthetic:\nminimal, elegant, expensive, technically sophisticated, cinematic, polished, restrained, premium AI launch campaign.\n\nAvoid:\n- excessive neon\n- cyberpunk look\n- too many thumbnails\n- busy collage layouts\n- random icons\n- fake UI elements\n- gradients inside every object\n- excessive lens flares\n- heavy shadows\n- 3D metallic typography\n- cartoon imagery\n- excessive text\n- logos other than the specified text\n- watermark\n- spelling errors\n- distorted letters\n- duplicated text",
26
  "steps": 8,
27
  "size_label": "3:2 · 2496×1664 (2048² area)",
 
35
  null
36
  ],
37
  "result": "t2i_capybara.png",
38
+ "info": "Pre-rendered example · seed `0` · 1248×832 · 6 steps · v0.2.1 LoRA r256 · prompt enhanced",
39
  "used_prompt": "The image is a close-up realistic photograph of a soaking wet capybara taking shelter under a large banana leaf in a rainy jungle, the background a lush green canopy dappled with falling raindrops and misty humidity. In the upper-left corner, the edge of the banana leaf curves downward, its broad surface glistening with rainwater and casting a soft shadow over the capybara’s back. Across the top third, rain streaks blur the dense foliage behind, with shafts of diffused light filtering through the canopy. In the center of the frame, the capybara’s rounded body is pressed against the leaf’s underside, its fur matted and dark with water, clinging to its skin in clumps; its small ears are flattened against its head, and its eyes are half-closed, glistening with moisture. On the right side, a second banana leaf peeks into the frame, its veins visible and slightly curled at the edges, partially obscuring the background. Behind the capybara’s left shoulder, a cluster of broad, wet leaves hangs low, their edges blurred by the shallow depth of field. Along the lower edge, the ground is a muddy brown, slick with rainwater and scattered with fallen leaves and twigs, reflecting the dim light. The lighting is soft, diffused daylight filtered through the jungle canopy, casting gentle highlights on the capybara’s wet fur and the water droplets clinging to the banana leaf, with faint shadows beneath the animal and along the leaf’s underside. The overall composition feels intimate and sheltered, emphasizing the capybara’s vulnerability and the quiet resilience of nature in the downpour.",
40
  "steps": 6,
41
  "size_label": "3:2 · 1248×832 (1024² area)",
 
49
  null
50
  ],
51
  "result": "t2i_diorama.png",
52
+ "info": "Pre-rendered example · seed `0` · 2368×1760 · 6 steps · v0.2.1 LoRA r256 · prompt enhanced",
53
  "used_prompt": "The image is a 4:3 isometric miniature diorama of a mountain research station, rendered in high detail with soft studio lighting and a tilt-shift effect. The scene is set on a rugged, snow-dusted mountain slope, with the research station occupying the central foreground, its modular buildings and domed observatory nestled into the terrain. To the left, a steep incline rises into a misty valley, while the right side features a rocky outcrop with a small antenna array. Behind the station, a jagged mountain ridge stretches across the background under a pale blue sky with soft, diffused clouds. The station’s structures are detailed with weathered metal siding, glass windows reflecting the ambient light, and exposed piping and cables. A small helipad with a landing pad marker is visible in the lower-left corner, and a narrow service road winds up the slope toward the main entrance. The lighting is soft and even, emanating from an unseen source above and slightly to the left, casting gentle shadows that emphasize the three-dimensional depth of the diorama. The overall composition feels immersive and meticulously crafted, with a balanced arrangement of architectural elements and natural terrain, evoking a sense of isolation and scientific exploration in a remote alpine environment.",
54
  "steps": 6,
55
  "size_label": "4:3 · 2368×1760 (2048² area)",
 
63
  null
64
  ],
65
  "result": "edit_sketch.png",
66
+ "info": "Pre-rendered example · seed `3` · 1024×1024 · 6 steps · v0.2.1 LoRA r256 · prompt enhanced",
67
  "used_prompt": "Convert the image to a pencil sketch on textured paper, preserving the exact composition, framing, and spatial arrangement of all elements including the rain-streaked window, the interior bookstore shelves, the figure inside, and the \"OPEN LATE\" sign in the foreground, while rendering the entire scene in a hand-drawn, sketchy style with visible pencil strokes and paper texture.",
68
  "steps": 6,
69
  "size_label": "Auto · match the last reference (1024² area)",
 
77
  null
78
  ],
79
  "result": "edit_2ref_cat.png",
80
+ "info": "Pre-rendered example · seed `42` · 832×1248 · 6 steps · v0.2.1 LoRA r256 · prompt enhanced",
81
  "used_prompt": "Compose a new image where the woman from <image1>, wearing her full traditional outfit with skull makeup and holding a smoke bomb, is now holding the fluffy cat from <image2> in her arms, positioned in the same spot as in the original image, with the ornate wooden gate and surrounding smoke from <image1> remaining unchanged as the background.",
82
  "steps": 6,
83
  "size_label": "Auto · match the last reference (1024² area)",
examples/t2i_capybara.png CHANGED

Git LFS Details

  • SHA256: 64d167cdcf2efdd523c42de494f7f5ecf74034aaf46fce0ef7d84de7c096a9b6
  • Pointer size: 132 Bytes
  • Size of remote file: 2.09 MB

Git LFS Details

  • SHA256: 424160187cf9cf97b021eb1d25aaefdd78bd48e9e07fe8c52f287985676930f4
  • Pointer size: 132 Bytes
  • Size of remote file: 2.09 MB
examples/t2i_diorama.png CHANGED

Git LFS Details

  • SHA256: 2010ee9ccb72cc02e24fb420246c589db4b4db3ab513131cb5d3b76d08e165b4
  • Pointer size: 132 Bytes
  • Size of remote file: 7.73 MB

Git LFS Details

  • SHA256: 00db3cb68fa023ee504c79442845b6bb9997cee9578d02b3c964d24f6f5799c9
  • Pointer size: 132 Bytes
  • Size of remote file: 7.75 MB
examples/t2i_launch.png CHANGED

Git LFS Details

  • SHA256: 1114e22b292914d036adae68f0006cd169d5363262b2017bae73875a0f033f96
  • Pointer size: 132 Bytes
  • Size of remote file: 5.15 MB

Git LFS Details

  • SHA256: 94246a34802366a17bfd7991f7583ab9efa1784f6181f0ecc14403916b50be01
  • Pointer size: 132 Bytes
  • Size of remote file: 5.15 MB