Spaces:
Running on Zero
Running on Zero
Comparison tab (33 Qwen Space examples, turbo 6 steps vs base 40 steps, image slider); 5 examples from the Qwen Space; header: very competitive with the base (part 2)
Browse files- .gitattributes +30 -0
- compare.py +141 -0
- compare/img/case25/turbo_t.webp +0 -0
- compare/img/en_1/base.webp +3 -0
- compare/img/en_1/base_t.webp +0 -0
- compare/img/en_1/turbo.webp +3 -0
- compare/img/en_1/turbo_t.webp +0 -0
- compare/img/en_2/base.webp +3 -0
- compare/img/en_2/base_t.webp +0 -0
- compare/img/en_2/turbo.webp +3 -0
- compare/img/en_2/turbo_t.webp +0 -0
- compare/img/en_3/base.webp +3 -0
- compare/img/en_3/base_t.webp +3 -0
- compare/img/en_3/turbo.webp +3 -0
- compare/img/en_3/turbo_t.webp +3 -0
- compare/img/en_4/base.webp +3 -0
- compare/img/en_4/base_t.webp +3 -0
- compare/img/en_4/turbo.webp +3 -0
- compare/img/en_4/turbo_t.webp +3 -0
- compare/img/zh_1/base.webp +3 -0
- compare/img/zh_1/base_t.webp +3 -0
- compare/img/zh_1/turbo.webp +3 -0
- compare/img/zh_1/turbo_t.webp +3 -0
- compare/img/zh_2/base.webp +3 -0
- compare/img/zh_2/base_t.webp +3 -0
- compare/img/zh_2/turbo.webp +3 -0
- compare/img/zh_2/turbo_t.webp +3 -0
- compare/img/zh_7/base.webp +3 -0
- compare/img/zh_7/base_t.webp +0 -0
- compare/img/zh_7/turbo.webp +3 -0
- compare/img/zh_7/turbo_t.webp +3 -0
- examples/manifest.json +70 -0
- examples/qwen_mask_edit.png +3 -0
- examples/qwen_mask_edit_in1.webp +3 -0
- examples/qwen_mask_edit_in2.webp +0 -0
- examples/qwen_perfume.png +3 -0
- examples/qwen_poster.png +3 -0
- examples/qwen_prompts.json +7 -0
- examples/qwen_rgba_bride.png +3 -0
- examples/qwen_storyboard.png +3 -0
- examples/qwen_storyboard_in1.webp +3 -0
.gitattributes
CHANGED
|
@@ -166,3 +166,33 @@ compare/img/case25/base.webp filter=lfs diff=lfs merge=lfs -text
|
|
| 166 |
compare/img/case25/in1.webp filter=lfs diff=lfs merge=lfs -text
|
| 167 |
compare/img/case25/ref.webp filter=lfs diff=lfs merge=lfs -text
|
| 168 |
compare/img/case25/turbo.webp filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 166 |
compare/img/case25/in1.webp filter=lfs diff=lfs merge=lfs -text
|
| 167 |
compare/img/case25/ref.webp filter=lfs diff=lfs merge=lfs -text
|
| 168 |
compare/img/case25/turbo.webp filter=lfs diff=lfs merge=lfs -text
|
| 169 |
+
compare/img/en_1/base.webp filter=lfs diff=lfs merge=lfs -text
|
| 170 |
+
compare/img/en_1/turbo.webp filter=lfs diff=lfs merge=lfs -text
|
| 171 |
+
compare/img/en_2/base.webp filter=lfs diff=lfs merge=lfs -text
|
| 172 |
+
compare/img/en_2/turbo.webp filter=lfs diff=lfs merge=lfs -text
|
| 173 |
+
compare/img/en_3/base.webp filter=lfs diff=lfs merge=lfs -text
|
| 174 |
+
compare/img/en_3/base_t.webp filter=lfs diff=lfs merge=lfs -text
|
| 175 |
+
compare/img/en_3/turbo.webp filter=lfs diff=lfs merge=lfs -text
|
| 176 |
+
compare/img/en_3/turbo_t.webp filter=lfs diff=lfs merge=lfs -text
|
| 177 |
+
compare/img/en_4/base.webp filter=lfs diff=lfs merge=lfs -text
|
| 178 |
+
compare/img/en_4/base_t.webp filter=lfs diff=lfs merge=lfs -text
|
| 179 |
+
compare/img/en_4/turbo.webp filter=lfs diff=lfs merge=lfs -text
|
| 180 |
+
compare/img/en_4/turbo_t.webp filter=lfs diff=lfs merge=lfs -text
|
| 181 |
+
compare/img/zh_1/base.webp filter=lfs diff=lfs merge=lfs -text
|
| 182 |
+
compare/img/zh_1/base_t.webp filter=lfs diff=lfs merge=lfs -text
|
| 183 |
+
compare/img/zh_1/turbo.webp filter=lfs diff=lfs merge=lfs -text
|
| 184 |
+
compare/img/zh_1/turbo_t.webp filter=lfs diff=lfs merge=lfs -text
|
| 185 |
+
compare/img/zh_2/base.webp filter=lfs diff=lfs merge=lfs -text
|
| 186 |
+
compare/img/zh_2/base_t.webp filter=lfs diff=lfs merge=lfs -text
|
| 187 |
+
compare/img/zh_2/turbo.webp filter=lfs diff=lfs merge=lfs -text
|
| 188 |
+
compare/img/zh_2/turbo_t.webp filter=lfs diff=lfs merge=lfs -text
|
| 189 |
+
compare/img/zh_7/base.webp filter=lfs diff=lfs merge=lfs -text
|
| 190 |
+
compare/img/zh_7/turbo.webp filter=lfs diff=lfs merge=lfs -text
|
| 191 |
+
compare/img/zh_7/turbo_t.webp filter=lfs diff=lfs merge=lfs -text
|
| 192 |
+
examples/qwen_mask_edit.png filter=lfs diff=lfs merge=lfs -text
|
| 193 |
+
examples/qwen_mask_edit_in1.webp filter=lfs diff=lfs merge=lfs -text
|
| 194 |
+
examples/qwen_perfume.png filter=lfs diff=lfs merge=lfs -text
|
| 195 |
+
examples/qwen_poster.png filter=lfs diff=lfs merge=lfs -text
|
| 196 |
+
examples/qwen_rgba_bride.png filter=lfs diff=lfs merge=lfs -text
|
| 197 |
+
examples/qwen_storyboard.png filter=lfs diff=lfs merge=lfs -text
|
| 198 |
+
examples/qwen_storyboard_in1.webp filter=lfs diff=lfs merge=lfs -text
|
compare.py
ADDED
|
@@ -0,0 +1,141 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""The demo's Comparison tab: 33 of the 37 examples of the official Qwen/Qwen-Image-2.1 Space, rendered with viggle-turbo
|
| 2 |
+
v0.2.1 in 6 steps (and in 8 steps on the 5 dense-text examples) and with the base model in 40 steps. Both get the same
|
| 3 |
+
prompt, input images and seed 42, with prompt enhancement off. No GPU: everything is pre-rendered. compare/ (WebP images
|
| 4 |
+
+ cases.json) is written by release/build_compare_space.py, and upload_space.sh builds it into the Space."""
|
| 5 |
+
import json
|
| 6 |
+
import statistics
|
| 7 |
+
from pathlib import Path
|
| 8 |
+
|
| 9 |
+
import gradio as gr
|
| 10 |
+
from PIL import Image
|
| 11 |
+
|
| 12 |
+
DIR = Path(__file__).parent / "compare"
|
| 13 |
+
CASES = json.loads((DIR / "cases.json").read_text(encoding="utf-8")) if (DIR / "cases.json").exists() else []
|
| 14 |
+
BY_ID = {case["id"]: case for case in CASES}
|
| 15 |
+
FIRST = "case01" # Character style infographic: a dense layout where we prefer the turbo's render
|
| 16 |
+
KINDS = {
|
| 17 |
+
"All": lambda case: True,
|
| 18 |
+
"Editing": lambda case: len(case["inputs"]) > 0,
|
| 19 |
+
"Multi-image editing": lambda case: len(case["inputs"]) > 1,
|
| 20 |
+
"Text-to-image": lambda case: not case["inputs"],
|
| 21 |
+
"Dense text (+8 steps)": lambda case: "turbo8" in case,
|
| 22 |
+
}
|
| 23 |
+
|
| 24 |
+
|
| 25 |
+
def arms(case):
|
| 26 |
+
"""(menu label, key) for every render of an example; the render time is in the label."""
|
| 27 |
+
menu = [(f"viggle-turbo · 6 steps · {case['turbo_s']:.1f} s", "turbo")]
|
| 28 |
+
if "turbo8" in case:
|
| 29 |
+
menu.append((f"viggle-turbo · 8 steps · {case['turbo8_s']:.1f} s", "turbo8"))
|
| 30 |
+
menu.append((f"Qwen-Image-2.1 base · 40 steps · {case['base_s']:.1f} s", "base"))
|
| 31 |
+
if case["reference"]:
|
| 32 |
+
menu.append(("Qwen API reference (bundled with Qwen's Space)", "reference"))
|
| 33 |
+
return menu
|
| 34 |
+
|
| 35 |
+
|
| 36 |
+
def image(case, key):
|
| 37 |
+
if key == "reference": # Qwen's API output has its own size; match the renders so the two slider halves line up
|
| 38 |
+
return Image.open(DIR / case["reference"]["src"]).resize(tuple(case["size"]), Image.LANCZOS)
|
| 39 |
+
return str(DIR / case[key]["src"])
|
| 40 |
+
|
| 41 |
+
|
| 42 |
+
def panels(case, left, right):
|
| 43 |
+
"""Slider pair, caption, input images and prompt for one example."""
|
| 44 |
+
names = dict((key, label) for label, key in arms(case))
|
| 45 |
+
width, height = case["size"]
|
| 46 |
+
title = case["title"] if case["title_zh"] == case["title"] else f"{case['title']} · {case['title_zh']}"
|
| 47 |
+
caption = (f"### {title}\n"
|
| 48 |
+
f"{width} × {height} · **left:** {names[left]} · **right:** {names[right]} · drag the handle to compare")
|
| 49 |
+
inputs = [(str(DIR / item["src"]), f"input {i + 1}") for i, item in enumerate(case["inputs"])]
|
| 50 |
+
return (image(case, left), image(case, right)), caption, inputs, case["prompt"]
|
| 51 |
+
|
| 52 |
+
|
| 53 |
+
def thumbs(ids):
|
| 54 |
+
return [(str(DIR / BY_ID[case_id]["turbo"]["thumb"]), BY_ID[case_id]["title"]) for case_id in ids]
|
| 55 |
+
|
| 56 |
+
|
| 57 |
+
def render():
|
| 58 |
+
"""Builds the tab inside the caller's gr.Blocks / gr.Tab context."""
|
| 59 |
+
edits = [case for case in CASES if case["inputs"]]
|
| 60 |
+
t2is = [case for case in CASES if not case["inputs"]]
|
| 61 |
+
median = lambda cases, key: statistics.median(case[key] for case in cases) # noqa: E731
|
| 62 |
+
speedup = sum(case["base_s"] for case in CASES) / sum(case["turbo_s"] for case in CASES)
|
| 63 |
+
gr.Markdown(
|
| 64 |
+
f"**viggle-turbo (6 steps) vs Qwen-Image-2.1 (40 steps)** on {len(CASES)} of the 37 examples of the official "
|
| 65 |
+
"[Qwen/Qwen-Image-2.1 Space](https://huggingface.co/spaces/Qwen/Qwen-Image-2.1). Both models get the same prompt, "
|
| 66 |
+
"input images and seed, with prompt enhancement off, one sample each and no seed picking. "
|
| 67 |
+
f"At **{speedup:.1f}× less time** the turbo is very competitive with the base model, and on some examples "
|
| 68 |
+
"(the character style infographic, shown first) we prefer its render; complicated edits can still fall short of the base. "
|
| 69 |
+
"Pick an example on the left, then drag the slider. The menus also offer the 8-step turbo on the dense-text examples "
|
| 70 |
+
"and Qwen's own API reference where the Space ships one.\n\n"
|
| 71 |
+
f"Median time, editing: **{median(edits, 'turbo_s'):.1f} s vs {median(edits, 'base_s'):.1f} s** · "
|
| 72 |
+
f"text-to-image (~4 MP): **{median(t2is, 'turbo_s'):.1f} s vs {median(t2is, 'base_s'):.1f} s** · "
|
| 73 |
+
f"all {len(CASES)} examples: **{speedup:.1f}× faster** (one NVIDIA B200, one pipeline call each)"
|
| 74 |
+
)
|
| 75 |
+
case = BY_ID[FIRST]
|
| 76 |
+
pair, caption, inputs, prompt = panels(case, "turbo", "base")
|
| 77 |
+
ids = gr.State([item["id"] for item in CASES])
|
| 78 |
+
with gr.Row():
|
| 79 |
+
with gr.Column(scale=1, min_width=280):
|
| 80 |
+
kind = gr.Radio([(f"{name} ({sum(map(test, CASES))})", name) for name, test in KINDS.items()], value="All", label="Show")
|
| 81 |
+
gallery = gr.Gallery(thumbs(ids.value), columns=3, height=760, object_fit="cover", allow_preview=False,
|
| 82 |
+
show_label=False, show_download_button=False, show_fullscreen_button=False)
|
| 83 |
+
with gr.Column(scale=3):
|
| 84 |
+
head = gr.Markdown(caption)
|
| 85 |
+
with gr.Row():
|
| 86 |
+
left = gr.Dropdown(arms(case), value="turbo", label="Left")
|
| 87 |
+
right = gr.Dropdown(arms(case), value="base", label="Right")
|
| 88 |
+
slider = gr.ImageSlider(pair, type="filepath", interactive=False, max_height=820, show_label=False)
|
| 89 |
+
with gr.Row():
|
| 90 |
+
shown_inputs = gr.Gallery(inputs, label="Input images", columns=3, height=220, visible=bool(inputs), scale=1)
|
| 91 |
+
shown_prompt = gr.Textbox(prompt, label="Prompt (sent as written, prompt enhancement off)", lines=8, max_lines=8,
|
| 92 |
+
interactive=False, show_copy_button=True, scale=2)
|
| 93 |
+
with gr.Accordion("How these were made", open=False):
|
| 94 |
+
gr.Markdown(
|
| 95 |
+
"- **Turbo:** Qwen-Image-2.1 + the viggle-turbo v0.2.1 LoRA (rank 256), 6 steps on sigmas "
|
| 96 |
+
"`[1, 0.9375, 0.875, 0.75, 0.5, 0.25]`, no classifier-free guidance, exactly as the Generate tab runs it.\n"
|
| 97 |
+
"- **Turbo, 8 steps** (the 5 dense-text examples): sigmas `[1, 0.9375, 0.875, 0.75, 0.625, 0.5, 0.25, 0.125]`. "
|
| 98 |
+
"At 6 steps small Latin text can print twice, like a double exposure (decided in the 0.75 → 0.5 step), and small "
|
| 99 |
+
"Chinese strokes can get colour blotches (the last 0.25 → 0 step); each added sigma splits one of the two. In our "
|
| 100 |
+
"OCR tests 8 steps raise word recall from 0.71 to 0.83 on the academic infographic's caption (16 seeds) and from "
|
| 101 |
+
"0.76 to 0.90 on the architecture board's Chinese labels (4 seeds); the base model scores 0.95 on both.\n"
|
| 102 |
+
"- **Base:** Qwen-Image-2.1, 40 steps, default scheduler, `true_cfg_scale=1.0` (the pipeline default).\n"
|
| 103 |
+
"- **Both:** diffusers `QwenImage21Pipeline` in bf16, seed 42. Input images are encoded at 1024² pixels; the output keeps "
|
| 104 |
+
"the aspect ratio of Qwen's reference at 2048² pixels for text-to-image and 1536² for editing, in multiples of 32.\n"
|
| 105 |
+
"- **Time:** one pipeline call on one NVIDIA B200 (text encoding, denoising and VAE decoding), after warm-up.\n"
|
| 106 |
+
"- **Left out (4 of 37):** three Chinese infographics and a subtitled storyboard whose short prompts leave the "
|
| 107 |
+
"on-image text to Qwen's prompt enhancement; without it both models fill them with made-up characters.\n"
|
| 108 |
+
"- **Qwen API reference:** the output bundled with the example in Qwen's Space, made with Qwen's API, which may "
|
| 109 |
+
"differ from the open weights. It is omitted for the 7 text-to-image examples Qwen generated with prompt enhancement on.\n\n"
|
| 110 |
+
"Prompts, input images and reference outputs come from the "
|
| 111 |
+
"[Qwen/Qwen-Image-2.1 Space](https://huggingface.co/spaces/Qwen/Qwen-Image-2.1) (Qwen Research License)."
|
| 112 |
+
)
|
| 113 |
+
|
| 114 |
+
def show(case_id, left_key, right_key, fixed="left"):
|
| 115 |
+
"""Keeps the chosen pair when the new example has it (turbo8 and reference are per-example), else turbo vs base.
|
| 116 |
+
If both menus land on the same render, the one the user just changed (`fixed`) wins and the other moves off it."""
|
| 117 |
+
case = BY_ID[case_id]
|
| 118 |
+
keys = [key for _, key in arms(case)]
|
| 119 |
+
left_key = left_key if left_key in keys else "turbo"
|
| 120 |
+
right_key = right_key if right_key in keys else "base"
|
| 121 |
+
if left_key == right_key:
|
| 122 |
+
other = next(key for key in ("base", "turbo") if key != left_key)
|
| 123 |
+
left_key, right_key = (left_key, other) if fixed == "left" else (other, right_key)
|
| 124 |
+
pair, caption, inputs, prompt = panels(case, left_key, right_key)
|
| 125 |
+
return (case_id, gr.update(choices=arms(case), value=left_key), gr.update(choices=arms(case), value=right_key),
|
| 126 |
+
pair, caption, gr.update(value=inputs, visible=bool(inputs)), prompt)
|
| 127 |
+
|
| 128 |
+
def pick(ids, left_key, right_key, evt: gr.SelectData):
|
| 129 |
+
return show(ids[evt.index], left_key, right_key)
|
| 130 |
+
|
| 131 |
+
def filter_cases(name):
|
| 132 |
+
ids = [case["id"] for case in CASES if KINDS[name](case)]
|
| 133 |
+
return ids, thumbs(ids)
|
| 134 |
+
|
| 135 |
+
current = gr.State(FIRST)
|
| 136 |
+
outputs = [current, left, right, slider, head, shown_inputs, shown_prompt]
|
| 137 |
+
kind.change(filter_cases, inputs=kind, outputs=[ids, gallery])
|
| 138 |
+
gallery.select(pick, inputs=[ids, left, right], outputs=outputs)
|
| 139 |
+
# .input, not .change: show() rewrites both menus, which must not fire another show()
|
| 140 |
+
left.input(show, inputs=[current, left, right], outputs=outputs)
|
| 141 |
+
right.input(lambda case_id, left_key, right_key: show(case_id, left_key, right_key, "right"), inputs=[current, left, right], outputs=outputs)
|
compare/img/case25/turbo_t.webp
ADDED
|
compare/img/en_1/base.webp
ADDED
|
Git LFS Details
|
compare/img/en_1/base_t.webp
ADDED
|
compare/img/en_1/turbo.webp
ADDED
|
Git LFS Details
|
compare/img/en_1/turbo_t.webp
ADDED
|
compare/img/en_2/base.webp
ADDED
|
Git LFS Details
|
compare/img/en_2/base_t.webp
ADDED
|
compare/img/en_2/turbo.webp
ADDED
|
Git LFS Details
|
compare/img/en_2/turbo_t.webp
ADDED
|
compare/img/en_3/base.webp
ADDED
|
Git LFS Details
|
compare/img/en_3/base_t.webp
ADDED
|
Git LFS Details
|
compare/img/en_3/turbo.webp
ADDED
|
Git LFS Details
|
compare/img/en_3/turbo_t.webp
ADDED
|
Git LFS Details
|
compare/img/en_4/base.webp
ADDED
|
Git LFS Details
|
compare/img/en_4/base_t.webp
ADDED
|
Git LFS Details
|
compare/img/en_4/turbo.webp
ADDED
|
Git LFS Details
|
compare/img/en_4/turbo_t.webp
ADDED
|
Git LFS Details
|
compare/img/zh_1/base.webp
ADDED
|
Git LFS Details
|
compare/img/zh_1/base_t.webp
ADDED
|
Git LFS Details
|
compare/img/zh_1/turbo.webp
ADDED
|
Git LFS Details
|
compare/img/zh_1/turbo_t.webp
ADDED
|
Git LFS Details
|
compare/img/zh_2/base.webp
ADDED
|
Git LFS Details
|
compare/img/zh_2/base_t.webp
ADDED
|
Git LFS Details
|
compare/img/zh_2/turbo.webp
ADDED
|
Git LFS Details
|
compare/img/zh_2/turbo_t.webp
ADDED
|
Git LFS Details
|
compare/img/zh_7/base.webp
ADDED
|
Git LFS Details
|
compare/img/zh_7/base_t.webp
ADDED
|
compare/img/zh_7/turbo.webp
ADDED
|
Git LFS Details
|
compare/img/zh_7/turbo_t.webp
ADDED
|
Git LFS Details
|
examples/manifest.json
CHANGED
|
@@ -27,6 +27,76 @@
|
|
| 27 |
"size_label": "3:2 · 2496×1664 (2048² area)",
|
| 28 |
"enhance": false
|
| 29 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 30 |
{
|
| 31 |
"prompt": "Soaking wet capybara taking shelter under a banana leaf in the rainy jungle, close up photo",
|
| 32 |
"refs": [
|
|
|
|
| 27 |
"size_label": "3:2 · 2496×1664 (2048² area)",
|
| 28 |
"enhance": false
|
| 29 |
},
|
| 30 |
+
{
|
| 31 |
+
"prompt": "Design a refined travel poster for a fictional night train. Render the headline exactly as \"THE MIDNIGHT EXPRESS\" and the subtitle \"A journey under the stars\". A silver train curves through dark blue mountains beneath a crescent moon. Art Deco geometry, ivory and gold lettering, clear typographic hierarchy, generous margins, print-ready composition.",
|
| 32 |
+
"refs": [
|
| 33 |
+
null,
|
| 34 |
+
null,
|
| 35 |
+
null
|
| 36 |
+
],
|
| 37 |
+
"result": "qwen_poster.png",
|
| 38 |
+
"info": "Pre-rendered example · seed `42` · 1664×2496 · 6 steps · v0.2.1 LoRA r256 · prompt sent as written (enhancement off)",
|
| 39 |
+
"used_prompt": "Design a refined travel poster for a fictional night train. Render the headline exactly as \"THE MIDNIGHT EXPRESS\" and the subtitle \"A journey under the stars\". A silver train curves through dark blue mountains beneath a crescent moon. Art Deco geometry, ivory and gold lettering, clear typographic hierarchy, generous margins, print-ready composition.",
|
| 40 |
+
"steps": 6,
|
| 41 |
+
"size_label": "2:3 · 1664×2496 (2048² area)",
|
| 42 |
+
"enhance": false
|
| 43 |
+
},
|
| 44 |
+
{
|
| 45 |
+
"prompt": "生成一张无缝的分镜插画,**现代真实偶像剧风格**(电影感写实摄影,柔和自然光与暖色调,浅景深虚化,韩系都市偶像剧质感,画面精致高级)。\n\n【版面(必须严格遵守)】\n- 恰好 6 个分镜格,**等宽**,排成**单独一横排(1 行 6 列)**;不是网格、不是两行,绝不要把任何一格再拆成上下两块。\n- 每格都是**竖长条矩形**(高明显大于宽,近似 9:16 竖条)。\n- 从左到右依次编号 1、2、3、4、5、6,每格左上角一个圆形数字徽标;**每个数字只出现一次,不重复、不跳号、不缺号。**\n- **每一格里都必须画出角色本人**在做该格的动作,角色占该格画面主体、完整清晰(半身或全身,约占该格高度 2/3 以上);**严禁出现只有背景、没有人物的空格。**\n- 6 格之间画风、光影氛围、人物一致。输出为一张扁平整图,6 格并排合并,不要输出多张分离文件。\n\n【角色】以输入参考图(含面部特写与正/侧/背三视全身立绘)为角色设计基准:现代年轻女性,棕色波浪长发带空气刘海,妆容清透,穿粉色方领修身长袖针织上衣、棕色灯芯绒过膝半裙配棕色皮带、棕色皮质机车靴。每一格里都是同一个人——面容、发型、服饰、身材比例保持一致,只改变姿势、表情、动作与所处场景;参考图仅定义外观,不要照搬其站姿。把角色自然地融入各自的真实场景中,透视比例正确、光影贴合。\n\n【逐格场景与内容】\n第1格:清晨阳光洒入的现代简约公寓,女孩站在落地窗边,双手捧着马克杯,望向窗外的城市天际线,神情恬静。\n第2格:白天繁华的都市街道人行道,两旁是时尚店铺橱窗和梧桐树,女孩背着单肩包边走边微微回头,浅浅微笑。\n第3格:文艺清新的咖啡馆靠窗卡座,女孩坐着低头看手机,嘴角带笑,窗外光线柔和,桌上放着一杯拿铁。\n第4格:暖色调的独立书店书架走廊,女孩踮脚从书架上抽出一本书,侧脸专注,光线温柔。\n第5格:傍晚金色晚霞下的城市人行天桥,女孩倚着栏杆回眸,背景是虚化的车流与高楼,暖光洒在脸上。\n第6格:夜晚霓虹灯光的街角,女孩仰头微笑看向前方,暖黄路灯与散景光斑环绕,都市夜景氛围。",
|
| 46 |
+
"refs": [
|
| 47 |
+
"qwen_storyboard_in1.webp",
|
| 48 |
+
null,
|
| 49 |
+
null
|
| 50 |
+
],
|
| 51 |
+
"result": "qwen_storyboard.png",
|
| 52 |
+
"info": "Pre-rendered example · seed `42` · 2048×1152 · 6 steps · v0.2.1 LoRA r256 · prompt sent as written (enhancement off)",
|
| 53 |
+
"used_prompt": "生成一张无缝的分镜插画,**现代真实偶像剧风格**(电影感写实摄影,柔和自然光与暖色调,浅景深虚化,韩系都市偶像剧质感,画面精致高级)。\n\n【版面(必须严格遵守)】\n- 恰好 6 个分镜格,**等宽**,排成**单独一横排(1 行 6 列)**;不是网格、不是两行,绝不要把任何一格再拆成上下两块。\n- 每格都是**竖长条矩形**(高明显大于宽,近似 9:16 竖条)。\n- 从左到右依次编号 1、2、3、4、5、6,每格左上角一个圆形数字徽标;**每个数字只出现一次,不重复、不跳号、不缺号。**\n- **每一格里都必须画出角色本人**在做该格的动作,角色占该格画面主体、完整清晰(半身或全身,约占该格高度 2/3 以上);**严禁出现只有背景、没有人物的空格。**\n- 6 格之间画风、光影氛围、人物一致。输出为一张扁平整图,6 格并排合并,不要输出多张分离文件。\n\n【角色】以输入参考图(含面部特写与正/侧/背三视全身立绘)为角色设计基准:现代年轻女性,棕色波浪长发带空气刘海,妆容清透,穿粉色方领修身长袖���织上衣、棕色灯芯绒过膝半裙配棕色皮带、棕色皮质机车靴。每一格里都是同一个人——面容、发型、服饰、身材比例保持一致,只改变姿势、表情、动作与所处场景;参考图仅定义外观,不要照搬其站姿。把角色自然地融入各自的真实场景中,透视比例正确、光影贴合。\n\n【逐格场景与内容】\n第1格:清晨阳光洒入的现代简约公寓,女孩站在落地窗边,双手捧着马克杯,望向窗外的城市天际线,神情恬静。\n第2格:白天繁华的都市街道人行道,两旁是时尚店铺橱窗和梧桐树,女孩背着单肩包边走边微微回头,浅浅微笑。\n第3格:文艺清新的咖啡馆靠窗卡座,女孩坐着低头看手机,嘴角带笑,窗外光线柔和,桌上放着一杯拿铁。\n第4格:暖色调的独立书店书架走廊,女孩踮脚从书架上抽出一本书,侧脸专注,光线温柔。\n第5格:傍晚金色晚霞下的城市人行天桥,女孩倚着栏杆回眸,背景是虚化的车流与高楼,暖光洒在脸上。\n第6格:夜晚霓虹灯光的街角,女孩仰头微笑看向前方,暖黄路灯与散景光斑环绕,都市夜景氛围。",
|
| 54 |
+
"steps": 6,
|
| 55 |
+
"size_label": "16:9 · 2048×1152 (1536² area)",
|
| 56 |
+
"enhance": false
|
| 57 |
+
},
|
| 58 |
+
{
|
| 59 |
+
"prompt": "Photograph a translucent emerald perfume bottle on pale limestone beside a shallow pool. Rippling sunlight reflects through the glass onto the stone. A single olive branch frames the upper left corner. Luxury product photography, realistic refraction, crisp bottle edges, soft shadows, uncluttered composition, no logo or text.",
|
| 60 |
+
"refs": [
|
| 61 |
+
null,
|
| 62 |
+
null,
|
| 63 |
+
null
|
| 64 |
+
],
|
| 65 |
+
"result": "qwen_perfume.png",
|
| 66 |
+
"info": "Pre-rendered example · seed `42` · 2496×1664 · 6 steps · v0.2.1 LoRA r256 · prompt sent as written (enhancement off)",
|
| 67 |
+
"used_prompt": "Photograph a translucent emerald perfume bottle on pale limestone beside a shallow pool. Rippling sunlight reflects through the glass onto the stone. A single olive branch frames the upper left corner. Luxury product photography, realistic refraction, crisp bottle edges, soft shadows, uncluttered composition, no logo or text.",
|
| 68 |
+
"steps": 6,
|
| 69 |
+
"size_label": "3:2 · 2496×1664 (2048² area)",
|
| 70 |
+
"enhance": false
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"prompt": "图中画圈标注的地方需要补上一名骑坐姿态的西部牛仔男子。他头戴棕色宽檐牛仔帽,蓄着浓密的络腮胡,脸侧向画面左方;上身穿棕色帆布夹克,里面搭配深蓝色牛仔衬衫,颈间系着浅棕色围巾;下身穿着带流苏的棕色皮质护腿,脚穿皮靴踩进马镫,一只手搭在鞍部附近,整体呈现自然的骑乘状态。",
|
| 74 |
+
"refs": [
|
| 75 |
+
"qwen_mask_edit_in1.webp",
|
| 76 |
+
"qwen_mask_edit_in2.webp",
|
| 77 |
+
null
|
| 78 |
+
],
|
| 79 |
+
"result": "qwen_mask_edit.png",
|
| 80 |
+
"info": "Pre-rendered example · seed `42` · 1248×1888 · 6 steps · v0.2.1 LoRA r256 · prompt sent as written (enhancement off)",
|
| 81 |
+
"used_prompt": "图中画圈标注的地方需要补上一名骑坐姿态的西部牛仔男子。他头戴棕色宽檐牛仔帽,蓄着浓密的络腮胡,脸侧向画面左方;上身穿棕色帆布夹克,里面搭配深蓝色牛仔衬衫,颈间系着浅棕色围巾;下身穿着带流苏的棕色皮质护腿,脚穿皮靴踩进马镫,一只手搭在鞍部附近,整体呈现自然的骑乘状态。",
|
| 82 |
+
"steps": 6,
|
| 83 |
+
"size_label": "2:3 · 1248×1888 (1536² area)",
|
| 84 |
+
"enhance": false
|
| 85 |
+
},
|
| 86 |
+
{
|
| 87 |
+
"prompt": "这是一张带有透明度的RGBA格式图像,呈现一位身着华丽婚纱的动漫风格女性角色。画面中无任何可识别的文字信息。角色为年轻女性,拥有白皙肌肤、红润脸颊与明亮红色眼眸,面带温柔微笑,表情甜美而幸福。她头戴精致银色皇冠与白色蕾丝头纱,金色长发自然垂落肩头,发丝柔顺且富有光泽。身穿多层荷叶边设计的粉白渐变抹胸婚纱,裙摆蓬松飘逸,层次丰富,材质轻盈如花瓣般展开;胸前手持一束由白色玫瑰与绿色叶片组成的捧花,花束饱满,绿叶点缀其间,增添清新感。双腿修长,穿着透明薄纱质感的过膝袜与饰有玫瑰图案的白色高跟鞋,姿态优雅站立,双手轻握捧花置于腹前。整体构图为全身正面视角,色彩柔和梦幻,以粉、白、绿为主色调,光影细腻,具有典型的日系二次元插画美学风格,适用于游戏立绘、视觉小说或数字艺术收藏等场景。 该图像具有alpha通道,背景是透明的。",
|
| 88 |
+
"refs": [
|
| 89 |
+
null,
|
| 90 |
+
null,
|
| 91 |
+
null
|
| 92 |
+
],
|
| 93 |
+
"result": "qwen_rgba_bride.png",
|
| 94 |
+
"info": "Pre-rendered example · seed `42` · 1664×2496 · 6 steps · v0.2.1 LoRA r256 · prompt sent as written (enhancement off)",
|
| 95 |
+
"used_prompt": "这是一张带有透明度的RGBA格式图像,呈现一位身着华丽婚纱的动漫风格女性角色。画面中无任何可识别的文字信息。角色为年轻女性,拥有白皙肌肤、红润脸颊与明亮红色眼眸,面带温柔微笑,表情甜美而幸福。她头戴精致银色皇冠与白色蕾丝头纱,金色长发自然垂落肩头,发丝柔顺且富有光泽。身穿多层荷叶边设计的粉白渐变抹胸婚纱,裙摆蓬松飘逸,层次丰富,材质轻盈如花瓣般展开;胸前手持一束由白色玫瑰与绿色叶片组成的捧花,花束饱满,绿叶点缀其间,增添清新感。双腿修长,穿着透明薄纱质感的过膝袜与饰有玫瑰图案的白色高跟鞋,姿态优雅站立,双手轻握捧花置于腹前。整体构图为全身正面视角,色彩柔和梦幻,以粉、白、绿为主色调,光影细腻,具有典型的日系二次元插画美学风格,适用于游戏立绘、视觉小说或数字艺术收藏等场景。 该图像具有alpha通道,背景是透明的。",
|
| 96 |
+
"steps": 6,
|
| 97 |
+
"size_label": "2:3 · 1664×2496 (2048² area)",
|
| 98 |
+
"enhance": false
|
| 99 |
+
},
|
| 100 |
{
|
| 101 |
"prompt": "Soaking wet capybara taking shelter under a banana leaf in the rainy jungle, close up photo",
|
| 102 |
"refs": [
|
examples/qwen_mask_edit.png
ADDED
|
Git LFS Details
|
examples/qwen_mask_edit_in1.webp
ADDED
|
Git LFS Details
|
examples/qwen_mask_edit_in2.webp
ADDED
|
examples/qwen_perfume.png
ADDED
|
Git LFS Details
|
examples/qwen_poster.png
ADDED
|
Git LFS Details
|
examples/qwen_prompts.json
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"qwen_poster": "Design a refined travel poster for a fictional night train. Render the headline exactly as \"THE MIDNIGHT EXPRESS\" and the subtitle \"A journey under the stars\". A silver train curves through dark blue mountains beneath a crescent moon. Art Deco geometry, ivory and gold lettering, clear typographic hierarchy, generous margins, print-ready composition.",
|
| 3 |
+
"qwen_storyboard": "生成一张无缝的分镜插画,**现代真实偶像剧风格**(电影感写实摄影,柔和自然光与暖色调,浅景深虚化,韩系都市偶像剧质感,画面精致高级)。\n\n【版面(必须严格遵守)】\n- 恰好 6 个分镜格,**等宽**,排成**单独一横排(1 行 6 列)**;不是网格、不是两行,绝不要把任何一格再拆成上下两块。\n- 每格都是**竖长条矩形**(高明显大于宽,近似 9:16 竖条)。\n- 从左到右依次编号 1、2、3、4、5、6,每格左上角一个圆形数字徽标;**每个数字只出现一次,不重复、不跳号、不缺号。**\n- **每一格里都必须画出角色本人**在做该格的动作,角色占该格画面主体、完整清晰(半身或全身,约占该格高度 2/3 以上);**严禁出现只有背景、没有人物的空格。**\n- 6 格之间画风、光影氛围、人物一致。输出为一张扁平整图,6 格并排合并,不要输出多张分离文件。\n\n【角色】以输入参考图(含面部特写与正/侧/背三视全身立绘)为角色设计基准:现代年轻女性,棕色波浪长发带空气刘海,妆容清透,穿粉色方领修身长袖针织上衣、棕色灯芯绒过膝半裙配棕色皮带、棕色皮质机车靴。每一格里都是同一个人——面容、发型、服饰、身材比例保持一致,只改变姿势、表情、动作与所处场景;参考图仅定义外观,不要照搬其站姿。把角色自然地融入各自的真实场景中,透视比例正确、光影贴合。\n\n【逐格场景与内容】\n第1格:清晨阳光洒入的现代简约公寓,女孩站在落地窗边,双手捧着马克杯,望向窗外的城市天际线,神情恬静。\n第2格:白天繁华的都市街道人行道,两旁是时尚店铺橱窗和梧桐树,女孩背着单肩包边走边微微回头,浅浅微笑。\n第3格:文艺清新的咖啡馆靠窗卡座,女孩坐着低头看手机,嘴角带笑,窗外光线柔和,桌上放着一杯拿铁。\n第4格:暖色调的独立书店书架走廊,女孩踮脚从书架上抽出一本书,侧脸专注,光线温柔。\n第5格:傍晚金色晚霞下的城市人行天桥,女孩倚着栏杆回眸,背景是虚化的车流与高楼,暖光洒在脸上。\n第6格:夜晚霓虹灯光的街角,女孩仰头微笑看向前方,暖黄路灯与散景光斑环绕,都市夜景氛围。",
|
| 4 |
+
"qwen_perfume": "Photograph a translucent emerald perfume bottle on pale limestone beside a shallow pool. Rippling sunlight reflects through the glass onto the stone. A single olive branch frames the upper left corner. Luxury product photography, realistic refraction, crisp bottle edges, soft shadows, uncluttered composition, no logo or text.",
|
| 5 |
+
"qwen_mask_edit": "图中画圈标注的地方需要补上一名骑坐姿态的西部牛仔男子。他头戴棕色宽檐牛仔帽,蓄着浓密的络腮胡,脸侧向画面左方;上身穿棕色帆布夹克,里面搭配深蓝色牛仔衬衫,颈间系着浅棕色围巾;下身穿着带流苏的棕色皮质护腿,脚穿皮靴踩进马镫,一只手搭在鞍部附近,整体呈现自然的骑乘状态。",
|
| 6 |
+
"qwen_rgba_bride": "这是一张带有透明度的RGBA格式图像,呈现一位身着华丽婚纱的动漫风格女性角色。画面中无任何可识别的文字信息。角色为年轻女性,拥有白皙肌肤、红润脸颊与明亮红色眼眸,面带温柔微笑,表情甜美而幸福。她头戴精致银色皇冠与白色蕾丝头纱,金色长发自然垂落肩头,发丝柔顺且富有光泽。身穿多层荷叶边设计的粉白渐变抹胸婚纱,裙摆蓬松飘逸,层次丰富,材质轻盈如花瓣般展开;胸前手持一束由白色玫瑰与绿色叶片组成的捧花,花束饱满,绿叶点缀其间,增添清新感。双腿修长,穿着透明薄纱质感的过膝袜与饰有玫瑰图案的白色高跟鞋,姿态优雅站立,双手轻握捧花置于腹前。整体构图为全身正面视角,色彩柔和梦幻,以粉、白、绿为主色调,光影细腻,具有典型的日系二次元插画美学风格,适用于游戏立绘、视觉小说或数字艺术收藏等场景。 该图像具有alpha通道,背景是透明的。"
|
| 7 |
+
}
|
examples/qwen_rgba_bride.png
ADDED
|
Git LFS Details
|
examples/qwen_storyboard.png
ADDED
|
Git LFS Details
|
examples/qwen_storyboard_in1.webp
ADDED
|
Git LFS Details
|