Spaces:
Running on Zero
Running on Zero
v0.2.1: xlarge card, no VAE tiling (tiled reference encode wrecked edits; 3-reference edit OOMed at 44.3 GiB on large)
Browse files- README.md +26 -37
- app.py +10 -18
- examples/edit_2ref_cat.png +2 -2
- examples/edit_3ref_klein.png +2 -2
- examples/edit_sketch.png +2 -2
- examples/t2i_diorama.png +2 -2
- examples/t2i_launch.png +2 -2
README.md
CHANGED
|
@@ -22,8 +22,8 @@ tags:
|
|
| 22 |
- dmd
|
| 23 |
# Hardware: this Space needs ZeroGPU, set in Space Settings. `suggested_hardware` is deliberately
|
| 24 |
# unset, because the only ZeroGPU value the Hub metadata accepts is the legacy `zero-a10g`, and a
|
| 25 |
-
# 24 GB A10G cannot hold this pipeline. `app.py` requests the
|
| 26 |
-
# @spaces.GPU(size="
|
| 27 |
# Secrets: set HF_TOKEN (read scope) while Viggle/Qwen-Image-2.1-viggle-turbo is private.
|
| 28 |
# `hf_oauth` is not needed: the app calls no user-scoped Hub API.
|
| 29 |
---
|
|
@@ -148,41 +148,30 @@ convenience layered on top of the recipe, not part of it: untick it when the wor
|
|
| 148 |
Weights are ~31.5 GiB resident in bf16 (13.25 transformer + 16.33 text encoder + 0.63 VAE + the r=256
|
| 149 |
adapter, 1.3 GiB, loaded in bf16). Measured on one B200 with the v0.2 LoRA at 5 steps, prompt enhancement on,
|
| 150 |
around a single `generate()` call including the prompt rewrite and the VAE encode/decode, with the allocator
|
| 151 |
-
cache dropped before each measurement (`release/measure_mem.py`
|
| 152 |
-
measurements at 4 steps plus the adapter:
|
| 153 |
-
|
| 154 |
-
|
| 155 |
-
|
| 156 |
-
|
|
| 157 |
-
|
|
| 158 |
-
| text-to-image |
|
| 159 |
-
| text-to-image | 2048×2048 |
|
| 160 |
-
| edit, 1 reference, auto | 1024×1024 |
|
| 161 |
-
| edit,
|
| 162 |
-
| edit,
|
| 163 |
-
| edit,
|
| 164 |
-
|
| 165 |
-
|
| 166 |
-
|
| 167 |
-
|
| 168 |
-
|
| 169 |
-
|
| 170 |
-
|
| 171 |
-
|
| 172 |
-
|
| 173 |
-
|
| 174 |
-
|
| 175 |
-
size="large")` at 1× quota and switches tiling on for every editing call and for the 1536²/2048²
|
| 176 |
-
text-to-image buckets; 1024²-bucket text-to-image stays untiled and bit-identical to the untiled
|
| 177 |
-
pipeline, tiled outputs differ from untiled ones by 1.6/255 mean absolute (measured at 2048×2048).
|
| 178 |
-
With that, the largest call on the menu peaks at 41.1 GiB allocated, ~3.5 GiB inside the card; the
|
| 179 |
-
*reserved* figures are what the caching allocator held on a card with room to spare - on the 44.7 GiB
|
| 180 |
-
card it releases cached blocks before failing, so the allocated peak is the binding number, and
|
| 181 |
-
`app.py` sets `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` to keep fragmentation down. Prompt
|
| 182 |
-
enhancement runs before the denoise with nothing else live (~2 GiB of KV cache over the resident
|
| 183 |
-
weights) and does not raise the peak. Wall-clock on the ZeroGPU card (half an RTX Pro 6000 Blackwell)
|
| 184 |
-
will be slower than the B200 numbers above, and the prompt rewrite adds 2-8 s on a B200; `duration=90`
|
| 185 |
-
covers both.
|
| 186 |
|
| 187 |
**Cold start** is dominated by the download: ~30.9 GiB of base safetensors from
|
| 188 |
`Qwen/Qwen-Image-2.1` (13.25 transformer + 16.33 text encoder + 1.26 VAE + 0.02 processor) plus
|
|
|
|
| 22 |
- dmd
|
| 23 |
# Hardware: this Space needs ZeroGPU, set in Space Settings. `suggested_hardware` is deliberately
|
| 24 |
# unset, because the only ZeroGPU value the Hub metadata accepts is the legacy `zero-a10g`, and a
|
| 25 |
+
# 24 GB A10G cannot hold this pipeline. `app.py` requests the 96 GB card with
|
| 26 |
+
# @spaces.GPU(size="xlarge") (no VAE tiling); see the Hardware section for the measurements behind that.
|
| 27 |
# Secrets: set HF_TOKEN (read scope) while Viggle/Qwen-Image-2.1-viggle-turbo is private.
|
| 28 |
# `hf_oauth` is not needed: the app calls no user-scoped Hub API.
|
| 29 |
---
|
|
|
|
| 148 |
Weights are ~31.5 GiB resident in bf16 (13.25 transformer + 16.33 text encoder + 0.63 VAE + the r=256
|
| 149 |
adapter, 1.3 GiB, loaded in bf16). Measured on one B200 with the v0.2 LoRA at 5 steps, prompt enhancement on,
|
| 150 |
around a single `generate()` call including the prompt rewrite and the VAE encode/decode, with the allocator
|
| 151 |
+
cache dropped before each measurement (`release/measure_mem.py`; the two 2048² rows are the v0.1 full-transformer
|
| 152 |
+
measurements at 4 steps plus the adapter). **No VAE tiling anywhere**: `vae.enable_tiling()` also tiles the
|
| 153 |
+
*encode* of the reference images (256 px tiles), and tiled reference latents wreck an edit - duplicated subjects,
|
| 154 |
+
wrong scale - while a tiled decode is not the validated pipeline either.
|
| 155 |
+
|
| 156 |
+
| call | output | peak allocated | peak reserved | wall clock (B200) |
|
| 157 |
+
|---|---|---|---|---|
|
| 158 |
+
| text-to-image | 1024×1024 | 38.24 GiB | 40.05 GiB | 0.8 s |
|
| 159 |
+
| text-to-image | 2048×2048 | ~58.3 GiB | ~66 GiB | 3.4 s (v0.1) |
|
| 160 |
+
| edit, 1 reference, auto | 1024×1024 | 40.13 GiB | 42.68 GiB | 1.1 s |
|
| 161 |
+
| edit, 2 references, auto | 832×1248 | 42.04 GiB | 44.71 GiB | 1.5 s |
|
| 162 |
+
| edit, 3 references, auto (first example) | 928×1152 | 44.29 GiB | 47.84 GiB | 2.9 s |
|
| 163 |
+
| edit, 3 references (largest) | 1760×1344 | ~52.8 GiB | ~59 GiB | 2.9 s (v0.1) |
|
| 164 |
+
|
| 165 |
+
ZeroGPU `large` is a 48 GB card with 44.7 GiB usable, so the three-reference edit at the 1024² bucket is
|
| 166 |
+
exactly the call that failed there (the CUDA OOM surfaces on ZeroGPU as `NVML_SUCCESS == r INTERNAL ASSERT
|
| 167 |
+
FAILED` from the caching allocator, because the container cannot query NVML for the OOM report), and the
|
| 168 |
+
2048² text-to-image and 1536² editing buckets never fit untiled. `app.py` therefore requests
|
| 169 |
+
`@spaces.GPU(duration=90, size="xlarge")`: the full 96 GB RTX Pro 6000 Blackwell, at **2× quota** (an
|
| 170 |
+
unauthenticated visitor's 2 min/day covers about three of the largest calls, a free account's 5 min about
|
| 171 |
+
seven; PRO 40 min). `app.py` also sets `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`. Prompt enhancement
|
| 172 |
+
runs before the denoise with nothing else live (~2 GiB of KV cache over the resident weights) and does not
|
| 173 |
+
raise the peak. Wall-clock on the ZeroGPU card will be slower than the B200 numbers above, and the prompt
|
| 174 |
+
rewrite adds 2-8 s on a B200; `duration=90` covers both.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 175 |
|
| 176 |
**Cold start** is dominated by the download: ~30.9 GiB of base safetensors from
|
| 177 |
`Qwen/Qwen-Image-2.1` (13.25 transformer + 16.33 text encoder + 1.26 VAE + 0.02 processor) plus
|
app.py
CHANGED
|
@@ -5,17 +5,18 @@
|
|
| 5 |
import importlib.util
|
| 6 |
import os
|
| 7 |
|
| 8 |
-
# Less allocator fragmentation
|
| 9 |
os.environ.setdefault("PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True")
|
| 10 |
|
| 11 |
if importlib.util.find_spec("spaces"):
|
| 12 |
import spaces
|
| 13 |
|
| 14 |
-
# size="
|
| 15 |
-
#
|
| 16 |
-
#
|
| 17 |
-
#
|
| 18 |
-
|
|
|
|
| 19 |
else:
|
| 20 |
gpu = lambda fn: fn # noqa: E731
|
| 21 |
|
|
@@ -78,7 +79,6 @@ T2I_CHOICES = [*_bucket(1024), *_bucket(2048)]
|
|
| 78 |
EDIT_CHOICES = [AUTO, *_bucket(1024), *_bucket(1536)]
|
| 79 |
# the (width, height) pairs the editing menu actually offers, in menu order
|
| 80 |
EDIT_DIMS = [SIZES[label] for label in EDIT_CHOICES if SIZES[label]]
|
| 81 |
-
TILE_ABOVE_PIXELS = 1_500_000 # between the 1024² bucket (≤ 1_056_768 px) and the 1536² bucket (≥ 2_359_296 px)
|
| 82 |
|
| 83 |
if STUDENT == "full":
|
| 84 |
transformer = QwenImage21Transformer2DModel.from_pretrained(
|
|
@@ -95,6 +95,7 @@ else:
|
|
| 95 |
pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_config(pipe.scheduler.config, shift_terminal=None)
|
| 96 |
pipe.to("cuda")
|
| 97 |
|
|
|
|
| 98 |
# transformers' Qwen3VLVisionPatchEmbed runs an nn.Conv3d whose kernel_size equals its stride, which is
|
| 99 |
# exactly a linear map over each flattened patch. cuDNN has no usable bf16 Conv3d kernel for this shape and
|
| 100 |
# falls back to one that costs ~30 s per reference image (measured; 356 ms even with memory to spare, versus
|
|
@@ -177,17 +178,8 @@ def generate(prompt, image_1, image_2, image_3, size_label, seed, randomize_seed
|
|
| 177 |
enhance_note = f" · enhance {time.perf_counter() - started:.1f}s"
|
| 178 |
if used_prompt == prompt:
|
| 179 |
enhance_note += " (rewrite failed to parse, original prompt used)"
|
| 180 |
-
# VAE tiling
|
| 181 |
-
#
|
| 182 |
-
# 44.3 GiB allocated for three references (42.0 for two), which is the OOM the first Space example hit
|
| 183 |
-
# (surfacing on ZeroGPU as "NVML_SUCCESS == r INTERNAL ASSERT FAILED" from the allocator). Tiled, the same
|
| 184 |
-
# call peaks at 40.3 GiB and the largest editing call (three references, 1760x1344) at 41.1 GiB, its
|
| 185 |
-
# denoising peak. Text-to-image at the 1024² bucket (38.2 GiB) stays untiled and bit-identical to the
|
| 186 |
-
# untiled pipeline; tiled decodes differ from untiled ones by 1.6/255 mean absolute (measured at 2048²).
|
| 187 |
-
if images or (width and width * height > TILE_ABOVE_PIXELS):
|
| 188 |
-
pipe.vae.enable_tiling()
|
| 189 |
-
else:
|
| 190 |
-
pipe.vae.disable_tiling()
|
| 191 |
generator = torch.Generator(device="cuda").manual_seed(seed)
|
| 192 |
started = time.perf_counter()
|
| 193 |
result = pipe(
|
|
|
|
| 5 |
import importlib.util
|
| 6 |
import os
|
| 7 |
|
| 8 |
+
# Less allocator fragmentation (untiled 2048^2 text-to-image reserves ~65 GiB of the 96 GB card).
|
| 9 |
os.environ.setdefault("PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True")
|
| 10 |
|
| 11 |
if importlib.util.find_spec("spaces"):
|
| 12 |
import spaces
|
| 13 |
|
| 14 |
+
# size="xlarge" (the full 96 GB RTX Pro 6000 Blackwell, 2x quota): the pipeline runs exactly as validated, with no VAE
|
| 15 |
+
# tiling anywhere - tiling the encode wrecks reference-conditioned edits (duplicated subjects, wrong scale), and on the
|
| 16 |
+
# 48 GB "large" card (44.7 GiB usable) the untiled calls do not fit: three-reference edit at a 1024^2-area output
|
| 17 |
+
# 44.3 GiB, 2048^2 text-to-image 57.0 GiB, three-reference edit at 1760x1344 51.5 GiB (measured, B200, bf16 - see README).
|
| 18 |
+
# duration=90 covers the prompt rewrite plus the largest call.
|
| 19 |
+
gpu = spaces.GPU(duration=90, size="xlarge")
|
| 20 |
else:
|
| 21 |
gpu = lambda fn: fn # noqa: E731
|
| 22 |
|
|
|
|
| 79 |
EDIT_CHOICES = [AUTO, *_bucket(1024), *_bucket(1536)]
|
| 80 |
# the (width, height) pairs the editing menu actually offers, in menu order
|
| 81 |
EDIT_DIMS = [SIZES[label] for label in EDIT_CHOICES if SIZES[label]]
|
|
|
|
| 82 |
|
| 83 |
if STUDENT == "full":
|
| 84 |
transformer = QwenImage21Transformer2DModel.from_pretrained(
|
|
|
|
| 95 |
pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_config(pipe.scheduler.config, shift_terminal=None)
|
| 96 |
pipe.to("cuda")
|
| 97 |
|
| 98 |
+
|
| 99 |
# transformers' Qwen3VLVisionPatchEmbed runs an nn.Conv3d whose kernel_size equals its stride, which is
|
| 100 |
# exactly a linear map over each flattened patch. cuDNN has no usable bf16 Conv3d kernel for this shape and
|
| 101 |
# falls back to one that costs ~30 s per reference image (measured; 356 ms even with memory to spare, versus
|
|
|
|
| 178 |
enhance_note = f" · enhance {time.perf_counter() - started:.1f}s"
|
| 179 |
if used_prompt == prompt:
|
| 180 |
enhance_note += " (rewrite failed to parse, original prompt used)"
|
| 181 |
+
# No VAE tiling anywhere (xlarge card): tiling the reference encode wrecks edits, and tiled decodes are not the
|
| 182 |
+
# validated pipeline either.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 183 |
generator = torch.Generator(device="cuda").manual_seed(seed)
|
| 184 |
started = time.perf_counter()
|
| 185 |
result = pipe(
|
examples/edit_2ref_cat.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
examples/edit_3ref_klein.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
examples/edit_sketch.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
examples/t2i_diorama.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
examples/t2i_launch.png
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|