yycc's picture
v0.3: 6 steps by default, new 9-step setting (7 turbo + 2 base-model steps)
bf38730 verified
|
Raw History Blame Contribute Delete
4.7 kB
---
title: Viggle Turbo for Qwen-Image-2.1
emoji: ⚡
colorFrom: indigo
colorTo: purple
sdk: gradio
sdk_version: 5.50.0
python_version: "3.12.12"
app_file: app.py
pinned: false
license: other
license_name: qwen-research
license_link: ./LICENSE
short_description: 6-step Qwen-Image-2.1, T2I + editing, vs-base comparison
models:
- Viggle/Qwen-Image-2.1-viggle-turbo
- Qwen/Qwen-Image-2.1
tags:
- text-to-image
- image-editing
- distillation
# Hardware: this Space needs ZeroGPU, set in Space Settings. `suggested_hardware` is deliberately
# unset, because the only ZeroGPU value the Hub metadata accepts is the legacy `zero-a10g`, and a
# 24 GB A10G cannot hold this pipeline. `app.py` requests the 96 GB card with
# @spaces.GPU(size="xlarge") (no VAE tiling).
# Secrets: set HF_TOKEN (read scope) while Viggle/Qwen-Image-2.1-viggle-turbo is private.
# `hf_oauth` is not needed: the app calls no user-scoped Hub API.
---
# Viggle Turbo v0.3 — 6-step Qwen-Image-2.1
A distilled [Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) that generates and edits images in
**6 steps** with **no classifier-free guidance**, about **5× faster** than the 40-step base model. On most prompts it
is hard to tell apart from the base model; small, dense text and complicated edits (multi-reference composition, face
swaps, identity-preserving edits) can still fall short of it.
**v0.3 (2026-09-29):** at 6 steps, less grain than v0.2.1 and a little softer on fine texture. We think 6 steps is
close to its capacity: every further gain we found cost something elsewhere. The new **9-step** setting runs 7 turbo
steps and lets the base model finish the last two: finer detail and small text right more often, at about 1.4–1.5× the
time of 6 steps.
Weights, numbers and ComfyUI workflows are on the
[model card](https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo).
**Generate:** leave **References** empty for text-to-image, or add one to six images for editing, composition or
style transfer; the prompt refers to them as image 1, image 2, … in the order they were added. **Steps** runs from 3
to 9, 6 is the default. The **Examples** rows are pre-rendered by this exact app; clicking one loads its prompt,
references, settings and seed. The reference photos `woman1/cat_window/bird.webp` and the three-reference prompt come
from the [black-forest-labs/flux-klein-9b-kv](https://huggingface.co/spaces/black-forest-labs/flux-klein-9b-kv) Space;
the `qwen_*` rows come from the [Qwen/Qwen-Image-2.1](https://huggingface.co/spaces/Qwen/Qwen-Image-2.1) Space.
**Comparison** (no GPU needed): 32 of the 37 examples of the Qwen/Qwen-Image-2.1 Space, pre-rendered with this model
and with the base model in 40 steps (same prompt, inputs and seed 42, prompt enhancement off, one sample each), in an
image slider, plus Qwen's own API output where their Space ships one.
## Prompt enhancement
**Enhance prompt**: **Auto** (default) rewrites prompts shorter than 30 tokens, **On** always rewrites, **Off** never
does. The rewrite is done by DeepSeek V4.1 Flash through OpenRouter, prompted with shortened versions of the system
prompts published with `Qwen/Qwen-Image-2.1-PE-T2I` / `-PE-I2I`. The rewritten prompt is shown under **Prompt sent to
the model**; if the call fails, the original prompt is used. The rewrite can add things the request did not ask for,
on-image text included, so switch it **Off** when the wording matters.
**Privacy.** When enhancement runs, the prompt and the reference images (downscaled to at most 0.5 MP) are sent to
OpenRouter and the provider it routes the request to; the request only allows providers that neither train on nor
retain the data. With **Off**, nothing leaves the Space.
## Running it
Environment: `LORA_FILE` (default the v0.3 r256 adapter; the v0.2.1 and v0.2 adapters in the model repo also work),
`STUDENT_REPO`, `BASE_MODEL_ID`, `HF_TOKEN` (read scope, needed only while the model repo is private),
`OPENROUTER_API_KEY` (prompt enhancement; without it every prompt is sent as written). The app needs ZeroGPU `xlarge`
(96 GB): the largest calls do not fit on the 48 GB card untiled, and tiling the VAE breaks edits.
## License
The base model is released under the [Qwen RESEARCH LICENSE AGREEMENT](./LICENSE) — **non-commercial: research or
evaluation purposes only**. This distilled derivative and this demo inherit that restriction. Commercial use requires
a separate license from Alibaba (`model-business@notice.qwencloud.com`). See `NOTICE` for the required attribution.
> Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi
> Laboratory Technology Co., Ltd. All Rights Reserved.