Spaces:
Running on Zero
Sync develop into main (#38)
Browse files* Check Upscale 4x by default and drop model name from its label (#19)
Simplifies the option to a plain checkbox toggle without exposing the
underlying model name in the UI.
Claude-Session: https://claude.ai/code/session_01W9GLmk1R3mFfb7otDoGiWX
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* Add detailed export/traceback logging to diagnose post-pipeline UI errors (#20)
The Space logs cut off right after the pipeline stages finish, with no
trace of what happens next in generate_video() (export_to_video and the
run_inference call boundary), even though the UI shows a generic Error.
Log frame metadata and full tracebacks at those points so the next
failure is diagnosable from the Space logs alone.
* Pre-download RIFE/upscale weights at startup instead of inside @spaces.GPU (#22)
The RIFE and 4xLSDIRCompact weight fetches are pure network/disk I/O with no
CUDA dependency, but were only downloaded lazily on first use inside the
metered @spaces.GPU allocation — costing a few seconds of ZeroGPU quota on
every cold container start. Move both fetches to app startup (alongside
load_pipeline()), mirroring how the main video model is already preloaded.
Claude-Session: https://claude.ai/code/session_011keM1seSWsBZqPHrzKTeCG
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* Batch frames and same-shaped tiles through the upscale model (#23)
upscale_frames processed one frame at a time, and within each frame
_tile_process called the model once per tile (batch=1 every time). With
SRVGGNetCompact this leaves most of the GPU idle per call, since a single
~256px tile is far under typical GPU parallel capacity.
Group frames of equal size into batches (FRAME_BATCH_SIZE) before tiling,
and within _tile_process group tile jobs by input-patch shape (interior
tiles are all the same size; only the last row/column differs, clipped
against the image edge) so each model() call is one batched forward pass
over same-shaped patches, chunked to MAX_TILE_BATCH to bound memory. The
per-tile crop/placement math is unchanged — verified index-for-index
equivalent to the prior per-tile loop via a numpy simulation with a
synthetic nonlinear "model" across several batch/ragged-size cases.
No GPU available in this environment to benchmark; needs validation on
the dev Space before merge.
* Increase upscale tile size from 256 to 512 (#24)
Larger tiles mean fewer, bigger model() calls per frame, which cuts
pad/crop overhead and Python/kernel-launch count and improves GPU
occupancy for this small SRVGGNetCompact network. 512 is still
conservative for a 16GB+ Spaces GPU (T4/L4/A10G).
No GPU available in this environment to benchmark memory headroom or
speedup; needs validation on the dev Space before merge.
* torch.compile the upscale model on CUDA (#26)
Wraps the model with torch.compile(dynamic=True) on the CUDA path for
kernel fusion. dynamic=True avoids a fresh recompile for every distinct
tile shape (interior tiles vs. the clipped last row/column) — there's
only a handful of those per video, but without it each new shape would
trigger its own recompile. Wrapped in try/except: if torch.compile fails
to initialize (e.g. missing/incompatible triton on the target GPU image),
the eager model is used as-is.
No GPU available in this environment to benchmark, and this is the
riskiest of the four upscale perf changes (compile warmup cost vs.
steady-state speedup, triton availability on the Space's GPU image) —
needs careful validation on the dev Space before merge.
* Log when torch.compile falls back to the eager upscale model (#31)
The try/except guard added in #26 silently swallowed torch.compile
initialization failures (e.g. missing/incompatible triton on the Space's
GPU image), so a fallback to the eager model would be invisible in the
logs — leaving no way to tell, from the deployed Space's console output,
whether the compiled path is actually active. Print the exception instead
of passing silently.
* Add liability disclaimer to the Space UI (#30)
Notes that outputs are provided "as is" and that users are solely
responsible for generated content and how it's used/shared.
Claude-Session: https://claude.ai/code/session_015woCw9matvyXW3Xnx2JwXG
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* Stop Inductor compilation progress from leaking into the Gradio UI (#33)
gr.Progress(track_tqdm=True) monkey-patches tqdm globally, so it was
picking up TorchInductor's own "Inductor Compilation" progress bars
(triggered by torch.compile on the upscale model, and by the quantized
transformer/text_encoder) regardless of their own disable flag. Drop
track_tqdm and report the generation stage's progress explicitly via
callback_on_step_end instead, matching the interpolation/upscale stages.
* Quantize transformer with weight-only fp8 instead of dynamic-activation fp8 (#37)
Float8DynamicActivationFloat8WeightConfig quantizes activations on every
F.linear call, which under eager execution (AOT deferred, #3) meant an
unplanned CUDA allocation per linear layer per step. That allocator churn
reliably crashed the prod ZeroGPU Space with an NVML assert in the caching
allocator (#35). Float8WeightOnlyConfig keeps the weight-memory savings
without any runtime activation-quantization allocations, trading away the
fp8x8 matmul speedup until AOT compilation (#36) makes the dynamic-activation
config safe again.
Claude-Session: https://claude.ai/code/session_01GRZRv1d2KinVjvMhJ4c7fF
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
- model/pipeline.py +13 -4
|
@@ -13,7 +13,7 @@ import torch
|
|
| 13 |
import torch._dynamo
|
| 14 |
from diffusers.pipelines.wan.pipeline_wan_i2v import WanImageToVideoPipeline
|
| 15 |
from PIL import Image
|
| 16 |
-
from torchao.quantization import
|
| 17 |
|
| 18 |
MODEL_ID = "thornmaze/WAMU_v3_WAN2.2_I2V_LIGHTNING"
|
| 19 |
|
|
@@ -41,12 +41,21 @@ def load_pipeline() -> WanImageToVideoPipeline:
|
|
| 41 |
pipe = WanImageToVideoPipeline.from_pretrained(MODEL_ID, torch_dtype=torch.bfloat16).to("cuda")
|
| 42 |
|
| 43 |
# Quantize to shrink the ~69GB bf16 pipe (drops per-call ZeroGPU init time and peak memory).
|
| 44 |
-
# Independent of AOT compilation (FR-6/C-2, deferred)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 45 |
quantize_(pipe.text_encoder, Int8WeightOnlyConfig())
|
| 46 |
torch._dynamo.reset()
|
| 47 |
-
quantize_(pipe.transformer,
|
| 48 |
torch._dynamo.reset()
|
| 49 |
-
quantize_(pipe.transformer_2,
|
| 50 |
torch._dynamo.reset()
|
| 51 |
|
| 52 |
return pipe
|
|
|
|
| 13 |
import torch._dynamo
|
| 14 |
from diffusers.pipelines.wan.pipeline_wan_i2v import WanImageToVideoPipeline
|
| 15 |
from PIL import Image
|
| 16 |
+
from torchao.quantization import Float8WeightOnlyConfig, Int8WeightOnlyConfig, quantize_
|
| 17 |
|
| 18 |
MODEL_ID = "thornmaze/WAMU_v3_WAN2.2_I2V_LIGHTNING"
|
| 19 |
|
|
|
|
| 41 |
pipe = WanImageToVideoPipeline.from_pretrained(MODEL_ID, torch_dtype=torch.bfloat16).to("cuda")
|
| 42 |
|
| 43 |
# Quantize to shrink the ~69GB bf16 pipe (drops per-call ZeroGPU init time and peak memory).
|
| 44 |
+
# Independent of AOT compilation (FR-6/C-2, deferred).
|
| 45 |
+
#
|
| 46 |
+
# Weight-only (not dynamic-activation) float8 for the transformers: the reference project uses
|
| 47 |
+
# Float8DynamicActivationFloat8WeightConfig, but only under spaces.aoti_load() (AOT-compiled),
|
| 48 |
+
# which pre-plans memory for the per-layer activation-quantize casts that config adds. Running
|
| 49 |
+
# that same config eagerly (as we do here, AOT deferred per #3) left every one of those casts as
|
| 50 |
+
# an unplanned CUDA allocation, which reliably crashed ZeroGPU with an NVML assert in the
|
| 51 |
+
# caching allocator (see #35). Weight-only quantization keeps the weight-memory savings without
|
| 52 |
+
# any runtime activation-quantization allocations, at the cost of the fp8x8 matmul speedup —
|
| 53 |
+
# revisit once AOT compilation (#36) makes the dynamic-activation config safe again.
|
| 54 |
quantize_(pipe.text_encoder, Int8WeightOnlyConfig())
|
| 55 |
torch._dynamo.reset()
|
| 56 |
+
quantize_(pipe.transformer, Float8WeightOnlyConfig())
|
| 57 |
torch._dynamo.reset()
|
| 58 |
+
quantize_(pipe.transformer_2, Float8WeightOnlyConfig())
|
| 59 |
torch._dynamo.reset()
|
| 60 |
|
| 61 |
return pipe
|