yycc commited on
Commit
0b7a73f
·
verified ·
1 Parent(s): a1e7d5b

v0.2.1: xlarge card, no VAE tiling (tiled reference encode wrecked edits; 3-reference edit OOMed at 44.3 GiB on large)

Browse files
README.md CHANGED
@@ -22,8 +22,8 @@ tags:
22
  - dmd
23
  # Hardware: this Space needs ZeroGPU, set in Space Settings. `suggested_hardware` is deliberately
24
  # unset, because the only ZeroGPU value the Hub metadata accepts is the legacy `zero-a10g`, and a
25
- # 24 GB A10G cannot hold this pipeline. `app.py` requests the 48 GB card with
26
- # @spaces.GPU(size="large"); see the Hardware section for the measurements behind that.
27
  # Secrets: set HF_TOKEN (read scope) while Viggle/Qwen-Image-2.1-viggle-turbo is private.
28
  # `hf_oauth` is not needed: the app calls no user-scoped Hub API.
29
  ---
@@ -148,41 +148,30 @@ convenience layered on top of the recipe, not part of it: untick it when the wor
148
  Weights are ~31.5 GiB resident in bf16 (13.25 transformer + 16.33 text encoder + 0.63 VAE + the r=256
149
  adapter, 1.3 GiB, loaded in bf16). Measured on one B200 with the v0.2 LoRA at 5 steps, prompt enhancement on,
150
  around a single `generate()` call including the prompt rewrite and the VAE encode/decode, with the allocator
151
- cache dropped before each measurement (`release/measure_mem.py`); the two 2048² rows are the v0.1 full-transformer
152
- measurements at 4 steps plus the adapter:
153
-
154
- | call | output | VAE tiling | peak allocated | peak reserved | wall clock (B200) |
155
- |---|---|---|---|---|---|
156
- | text-to-image | 1024×1024 | off | 38.24 GiB | 40.05 GiB | 0.8 s |
157
- | text-to-image | 1024×1024 | on | 32.39 GiB | 32.90 GiB | 1.1 s |
158
- | text-to-image | 2048×2048 | off | ~58.3 GiB | ~66 GiB | 3.4 s (v0.1) |
159
- | text-to-image | 2048×2048 | **on** | ~34.8 GiB | ~35.5 GiB | 4.7 s (v0.1) |
160
- | edit, 1 reference, auto | 1024×1024 | off | 40.13 GiB | 42.68 GiB | 1.1 s |
161
- | edit, 1 reference, auto | 1024×1024 | **on** | 34.91 GiB | 36.36 GiB | 1.8 s |
162
- | edit, 2 references, auto | 832×1248 | off | 42.04 GiB | 44.71 GiB | 1.5 s |
163
- | edit, 2 references, auto | 832×1248 | **on** | 37.52 GiB | 39.82 GiB | 2.3 s |
164
- | edit, 3 references, auto (first example) | 928×1152 | off | **44.29 GiB** | 47.84 GiB | 2.9 s |
165
- | edit, 3 references, auto (first example) | 928×1152 | **on** | 40.29 GiB | 43.58 GiB | 3.0 s |
166
- | edit, 3 references (largest) | 1760×1344 | off | 41.11 GiB | 45.14 GiB | 5.2 s |
167
- | edit, 3 references (largest) | 1760×1344 | **on** | **41.11 GiB** | 45.14 GiB | 4.8 s |
168
-
169
- The untiled VAE decode is the peak of every call at the 1024² bucket once references are resident, and
170
- of every larger text-to-image; the three-reference edit at 1760×1344 peaks in the denoising loop
171
- instead (41.1 GiB, the same tiled or not). ZeroGPU **`large`** is a 48 GB card with 44.7 GiB usable,
172
- so the untiled three-reference edit at 44.3 GiB is exactly the call that failed on the Space (the CUDA
173
- OOM surfaces there as `NVML_SUCCESS == r INTERNAL ASSERT FAILED` from the caching allocator, because the
174
- container cannot query NVML for the OOM report). `app.py` therefore requests `@spaces.GPU(duration=90,
175
- size="large")` at 1× quota and switches tiling on for every editing call and for the 1536²/2048²
176
- text-to-image buckets; 1024²-bucket text-to-image stays untiled and bit-identical to the untiled
177
- pipeline, tiled outputs differ from untiled ones by 1.6/255 mean absolute (measured at 2048×2048).
178
- With that, the largest call on the menu peaks at 41.1 GiB allocated, ~3.5 GiB inside the card; the
179
- *reserved* figures are what the caching allocator held on a card with room to spare - on the 44.7 GiB
180
- card it releases cached blocks before failing, so the allocated peak is the binding number, and
181
- `app.py` sets `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` to keep fragmentation down. Prompt
182
- enhancement runs before the denoise with nothing else live (~2 GiB of KV cache over the resident
183
- weights) and does not raise the peak. Wall-clock on the ZeroGPU card (half an RTX Pro 6000 Blackwell)
184
- will be slower than the B200 numbers above, and the prompt rewrite adds 2-8 s on a B200; `duration=90`
185
- covers both.
186
 
187
  **Cold start** is dominated by the download: ~30.9 GiB of base safetensors from
188
  `Qwen/Qwen-Image-2.1` (13.25 transformer + 16.33 text encoder + 1.26 VAE + 0.02 processor) plus
 
22
  - dmd
23
  # Hardware: this Space needs ZeroGPU, set in Space Settings. `suggested_hardware` is deliberately
24
  # unset, because the only ZeroGPU value the Hub metadata accepts is the legacy `zero-a10g`, and a
25
+ # 24 GB A10G cannot hold this pipeline. `app.py` requests the 96 GB card with
26
+ # @spaces.GPU(size="xlarge") (no VAE tiling); see the Hardware section for the measurements behind that.
27
  # Secrets: set HF_TOKEN (read scope) while Viggle/Qwen-Image-2.1-viggle-turbo is private.
28
  # `hf_oauth` is not needed: the app calls no user-scoped Hub API.
29
  ---
 
148
  Weights are ~31.5 GiB resident in bf16 (13.25 transformer + 16.33 text encoder + 0.63 VAE + the r=256
149
  adapter, 1.3 GiB, loaded in bf16). Measured on one B200 with the v0.2 LoRA at 5 steps, prompt enhancement on,
150
  around a single `generate()` call including the prompt rewrite and the VAE encode/decode, with the allocator
151
+ cache dropped before each measurement (`release/measure_mem.py`; the two 2048² rows are the v0.1 full-transformer
152
+ measurements at 4 steps plus the adapter). **No VAE tiling anywhere**: `vae.enable_tiling()` also tiles the
153
+ *encode* of the reference images (256 px tiles), and tiled reference latents wreck an edit - duplicated subjects,
154
+ wrong scale - while a tiled decode is not the validated pipeline either.
155
+
156
+ | call | output | peak allocated | peak reserved | wall clock (B200) |
157
+ |---|---|---|---|---|
158
+ | text-to-image | 1024×1024 | 38.24 GiB | 40.05 GiB | 0.8 s |
159
+ | text-to-image | 2048×2048 | ~58.3 GiB | ~66 GiB | 3.4 s (v0.1) |
160
+ | edit, 1 reference, auto | 1024×1024 | 40.13 GiB | 42.68 GiB | 1.1 s |
161
+ | edit, 2 references, auto | 832×1248 | 42.04 GiB | 44.71 GiB | 1.5 s |
162
+ | edit, 3 references, auto (first example) | 928×1152 | 44.29 GiB | 47.84 GiB | 2.9 s |
163
+ | edit, 3 references (largest) | 1760×1344 | ~52.8 GiB | ~59 GiB | 2.9 s (v0.1) |
164
+
165
+ ZeroGPU `large` is a 48 GB card with 44.7 GiB usable, so the three-reference edit at the 1024² bucket is
166
+ exactly the call that failed there (the CUDA OOM surfaces on ZeroGPU as `NVML_SUCCESS == r INTERNAL ASSERT
167
+ FAILED` from the caching allocator, because the container cannot query NVML for the OOM report), and the
168
+ 2048² text-to-image and 1536² editing buckets never fit untiled. `app.py` therefore requests
169
+ `@spaces.GPU(duration=90, size="xlarge")`: the full 96 GB RTX Pro 6000 Blackwell, at **2× quota** (an
170
+ unauthenticated visitor's 2 min/day covers about three of the largest calls, a free account's 5 min about
171
+ seven; PRO 40 min). `app.py` also sets `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`. Prompt enhancement
172
+ runs before the denoise with nothing else live (~2 GiB of KV cache over the resident weights) and does not
173
+ raise the peak. Wall-clock on the ZeroGPU card will be slower than the B200 numbers above, and the prompt
174
+ rewrite adds 2-8 s on a B200; `duration=90` covers both.
 
 
 
 
 
 
 
 
 
 
 
175
 
176
  **Cold start** is dominated by the download: ~30.9 GiB of base safetensors from
177
  `Qwen/Qwen-Image-2.1` (13.25 transformer + 16.33 text encoder + 1.26 VAE + 0.02 processor) plus
app.py CHANGED
@@ -5,17 +5,18 @@
5
  import importlib.util
6
  import os
7
 
8
- # Less allocator fragmentation: the largest calls run within ~3.5 GiB of the card (see generate()).
9
  os.environ.setdefault("PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True")
10
 
11
  if importlib.util.find_spec("spaces"):
12
  import spaces
13
 
14
- # size="large" (48 GB, 1x quota) is enough once the VAE decode is tiled above the 1024^2 bucket and for
15
- # every editing call: the largest call (three-reference edit at 1760x1344) then peaks at 41.1 GiB
16
- # allocated, versus 57.0 / 64.7 GiB for an untiled 2048^2 text-to-image (measured, B200, bf16 - see README).
17
- # duration=90 covers the prompt rewrite (4-11 s on a B200, slower here) plus the largest call.
18
- gpu = spaces.GPU(duration=90, size="large")
 
19
  else:
20
  gpu = lambda fn: fn # noqa: E731
21
 
@@ -78,7 +79,6 @@ T2I_CHOICES = [*_bucket(1024), *_bucket(2048)]
78
  EDIT_CHOICES = [AUTO, *_bucket(1024), *_bucket(1536)]
79
  # the (width, height) pairs the editing menu actually offers, in menu order
80
  EDIT_DIMS = [SIZES[label] for label in EDIT_CHOICES if SIZES[label]]
81
- TILE_ABOVE_PIXELS = 1_500_000 # between the 1024² bucket (≤ 1_056_768 px) and the 1536² bucket (≥ 2_359_296 px)
82
 
83
  if STUDENT == "full":
84
  transformer = QwenImage21Transformer2DModel.from_pretrained(
@@ -95,6 +95,7 @@ else:
95
  pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_config(pipe.scheduler.config, shift_terminal=None)
96
  pipe.to("cuda")
97
 
 
98
  # transformers' Qwen3VLVisionPatchEmbed runs an nn.Conv3d whose kernel_size equals its stride, which is
99
  # exactly a linear map over each flattened patch. cuDNN has no usable bf16 Conv3d kernel for this shape and
100
  # falls back to one that costs ~30 s per reference image (measured; 356 ms even with memory to spare, versus
@@ -177,17 +178,8 @@ def generate(prompt, image_1, image_2, image_3, size_label, seed, randomize_seed
177
  enhance_note = f" · enhance {time.perf_counter() - started:.1f}s"
178
  if used_prompt == prompt:
179
  enhance_note += " (rewrite failed to parse, original prompt used)"
180
- # VAE tiling wherever the untiled decode would not fit the 48 GB (44.7 GiB) card: the 1536²/2048² buckets,
181
- # and every editing call - with references resident the untiled decode of a 1024²-area output peaks at
182
- # 44.3 GiB allocated for three references (42.0 for two), which is the OOM the first Space example hit
183
- # (surfacing on ZeroGPU as "NVML_SUCCESS == r INTERNAL ASSERT FAILED" from the allocator). Tiled, the same
184
- # call peaks at 40.3 GiB and the largest editing call (three references, 1760x1344) at 41.1 GiB, its
185
- # denoising peak. Text-to-image at the 1024² bucket (38.2 GiB) stays untiled and bit-identical to the
186
- # untiled pipeline; tiled decodes differ from untiled ones by 1.6/255 mean absolute (measured at 2048²).
187
- if images or (width and width * height > TILE_ABOVE_PIXELS):
188
- pipe.vae.enable_tiling()
189
- else:
190
- pipe.vae.disable_tiling()
191
  generator = torch.Generator(device="cuda").manual_seed(seed)
192
  started = time.perf_counter()
193
  result = pipe(
 
5
  import importlib.util
6
  import os
7
 
8
+ # Less allocator fragmentation (untiled 2048^2 text-to-image reserves ~65 GiB of the 96 GB card).
9
  os.environ.setdefault("PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True")
10
 
11
  if importlib.util.find_spec("spaces"):
12
  import spaces
13
 
14
+ # size="xlarge" (the full 96 GB RTX Pro 6000 Blackwell, 2x quota): the pipeline runs exactly as validated, with no VAE
15
+ # tiling anywhere - tiling the encode wrecks reference-conditioned edits (duplicated subjects, wrong scale), and on the
16
+ # 48 GB "large" card (44.7 GiB usable) the untiled calls do not fit: three-reference edit at a 1024^2-area output
17
+ # 44.3 GiB, 2048^2 text-to-image 57.0 GiB, three-reference edit at 1760x1344 51.5 GiB (measured, B200, bf16 - see README).
18
+ # duration=90 covers the prompt rewrite plus the largest call.
19
+ gpu = spaces.GPU(duration=90, size="xlarge")
20
  else:
21
  gpu = lambda fn: fn # noqa: E731
22
 
 
79
  EDIT_CHOICES = [AUTO, *_bucket(1024), *_bucket(1536)]
80
  # the (width, height) pairs the editing menu actually offers, in menu order
81
  EDIT_DIMS = [SIZES[label] for label in EDIT_CHOICES if SIZES[label]]
 
82
 
83
  if STUDENT == "full":
84
  transformer = QwenImage21Transformer2DModel.from_pretrained(
 
95
  pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_config(pipe.scheduler.config, shift_terminal=None)
96
  pipe.to("cuda")
97
 
98
+
99
  # transformers' Qwen3VLVisionPatchEmbed runs an nn.Conv3d whose kernel_size equals its stride, which is
100
  # exactly a linear map over each flattened patch. cuDNN has no usable bf16 Conv3d kernel for this shape and
101
  # falls back to one that costs ~30 s per reference image (measured; 356 ms even with memory to spare, versus
 
178
  enhance_note = f" · enhance {time.perf_counter() - started:.1f}s"
179
  if used_prompt == prompt:
180
  enhance_note += " (rewrite failed to parse, original prompt used)"
181
+ # No VAE tiling anywhere (xlarge card): tiling the reference encode wrecks edits, and tiled decodes are not the
182
+ # validated pipeline either.
 
 
 
 
 
 
 
 
 
183
  generator = torch.Generator(device="cuda").manual_seed(seed)
184
  started = time.perf_counter()
185
  result = pipe(
examples/edit_2ref_cat.png CHANGED

Git LFS Details

  • SHA256: f95ed77823db89a02cde0ed74abc83acd70bd2c7d7c1ae7c33091f7715a3c400
  • Pointer size: 132 Bytes
  • Size of remote file: 2.24 MB

Git LFS Details

  • SHA256: 4b1a69345eb7dd8f3d60aeea999a539f8af69c8ae500409f327ffe31a84f1054
  • Pointer size: 132 Bytes
  • Size of remote file: 2.09 MB
examples/edit_3ref_klein.png CHANGED

Git LFS Details

  • SHA256: 8fd5ad38d86055402f095eee3247c37e49d1d76779a6959f05ea604c662f44f4
  • Pointer size: 132 Bytes
  • Size of remote file: 1.97 MB

Git LFS Details

  • SHA256: 98c7142d2d849a94a8ec4356a451cac56ca31b75366449f115898b4820254108
  • Pointer size: 132 Bytes
  • Size of remote file: 1.72 MB
examples/edit_sketch.png CHANGED

Git LFS Details

  • SHA256: 6091577f62665cbb25357fe5af884481ab27eed9a8d691d9463d0b8485772dfa
  • Pointer size: 132 Bytes
  • Size of remote file: 2.43 MB

Git LFS Details

  • SHA256: d52c768e83acd8e01b7ed38e3897d0eb2f65cfd33b9cc7f0ee635eb00d244a6b
  • Pointer size: 132 Bytes
  • Size of remote file: 2.41 MB
examples/t2i_diorama.png CHANGED

Git LFS Details

  • SHA256: c2f4412e8bb3edeabb480b5ac79fccfad87120923a6aac1278b8291b6446e131
  • Pointer size: 132 Bytes
  • Size of remote file: 7.66 MB

Git LFS Details

  • SHA256: 6afda62e4d97a21cd2062010ae411c45679185516fb457c539fe13bf34c1d0fe
  • Pointer size: 132 Bytes
  • Size of remote file: 7.79 MB
examples/t2i_launch.png CHANGED

Git LFS Details

  • SHA256: b132dbd35f40975af7d64cdea21c84e6dcdb6930d9400b49cb83cec63ab3925c
  • Pointer size: 132 Bytes
  • Size of remote file: 5.85 MB

Git LFS Details

  • SHA256: 742f2c27ff10391d93e8d24310e7c809965e8b6777e8bcb70629affaee4429f1
  • Pointer size: 132 Bytes
  • Size of remote file: 5.81 MB