someone-in-the-world Claude Sonnet 5 commited on
Commit
9f64382
·
unverified ·
1 Parent(s): 5ab12ca

Sync develop into main (#38)

Browse files

* Check Upscale 4x by default and drop model name from its label (#19)

Simplifies the option to a plain checkbox toggle without exposing the
underlying model name in the UI.


Claude-Session: https://claude.ai/code/session_01W9GLmk1R3mFfb7otDoGiWX

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* Add detailed export/traceback logging to diagnose post-pipeline UI errors (#20)

The Space logs cut off right after the pipeline stages finish, with no
trace of what happens next in generate_video() (export_to_video and the
run_inference call boundary), even though the UI shows a generic Error.
Log frame metadata and full tracebacks at those points so the next
failure is diagnosable from the Space logs alone.

* Pre-download RIFE/upscale weights at startup instead of inside @spaces.GPU (#22)

The RIFE and 4xLSDIRCompact weight fetches are pure network/disk I/O with no
CUDA dependency, but were only downloaded lazily on first use inside the
metered @spaces.GPU allocation — costing a few seconds of ZeroGPU quota on
every cold container start. Move both fetches to app startup (alongside
load_pipeline()), mirroring how the main video model is already preloaded.


Claude-Session: https://claude.ai/code/session_011keM1seSWsBZqPHrzKTeCG

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* Batch frames and same-shaped tiles through the upscale model (#23)

upscale_frames processed one frame at a time, and within each frame
_tile_process called the model once per tile (batch=1 every time). With
SRVGGNetCompact this leaves most of the GPU idle per call, since a single
~256px tile is far under typical GPU parallel capacity.

Group frames of equal size into batches (FRAME_BATCH_SIZE) before tiling,
and within _tile_process group tile jobs by input-patch shape (interior
tiles are all the same size; only the last row/column differs, clipped
against the image edge) so each model() call is one batched forward pass
over same-shaped patches, chunked to MAX_TILE_BATCH to bound memory. The
per-tile crop/placement math is unchanged — verified index-for-index
equivalent to the prior per-tile loop via a numpy simulation with a
synthetic nonlinear "model" across several batch/ragged-size cases.

No GPU available in this environment to benchmark; needs validation on
the dev Space before merge.

* Increase upscale tile size from 256 to 512 (#24)

Larger tiles mean fewer, bigger model() calls per frame, which cuts
pad/crop overhead and Python/kernel-launch count and improves GPU
occupancy for this small SRVGGNetCompact network. 512 is still
conservative for a 16GB+ Spaces GPU (T4/L4/A10G).

No GPU available in this environment to benchmark memory headroom or
speedup; needs validation on the dev Space before merge.

* torch.compile the upscale model on CUDA (#26)

Wraps the model with torch.compile(dynamic=True) on the CUDA path for
kernel fusion. dynamic=True avoids a fresh recompile for every distinct
tile shape (interior tiles vs. the clipped last row/column) — there's
only a handful of those per video, but without it each new shape would
trigger its own recompile. Wrapped in try/except: if torch.compile fails
to initialize (e.g. missing/incompatible triton on the target GPU image),
the eager model is used as-is.

No GPU available in this environment to benchmark, and this is the
riskiest of the four upscale perf changes (compile warmup cost vs.
steady-state speedup, triton availability on the Space's GPU image) —
needs careful validation on the dev Space before merge.

* Log when torch.compile falls back to the eager upscale model (#31)

The try/except guard added in #26 silently swallowed torch.compile
initialization failures (e.g. missing/incompatible triton on the Space's
GPU image), so a fallback to the eager model would be invisible in the
logs — leaving no way to tell, from the deployed Space's console output,
whether the compiled path is actually active. Print the exception instead
of passing silently.

* Add liability disclaimer to the Space UI (#30)

Notes that outputs are provided "as is" and that users are solely
responsible for generated content and how it's used/shared.


Claude-Session: https://claude.ai/code/session_015woCw9matvyXW3Xnx2JwXG

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* Stop Inductor compilation progress from leaking into the Gradio UI (#33)

gr.Progress(track_tqdm=True) monkey-patches tqdm globally, so it was
picking up TorchInductor's own "Inductor Compilation" progress bars
(triggered by torch.compile on the upscale model, and by the quantized
transformer/text_encoder) regardless of their own disable flag. Drop
track_tqdm and report the generation stage's progress explicitly via
callback_on_step_end instead, matching the interpolation/upscale stages.

* Quantize transformer with weight-only fp8 instead of dynamic-activation fp8 (#37)

Float8DynamicActivationFloat8WeightConfig quantizes activations on every
F.linear call, which under eager execution (AOT deferred, #3) meant an
unplanned CUDA allocation per linear layer per step. That allocator churn
reliably crashed the prod ZeroGPU Space with an NVML assert in the caching
allocator (#35). Float8WeightOnlyConfig keeps the weight-memory savings
without any runtime activation-quantization allocations, trading away the
fp8x8 matmul speedup until AOT compilation (#36) makes the dynamic-activation
config safe again.


Claude-Session: https://claude.ai/code/session_01GRZRv1d2KinVjvMhJ4c7fF

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

Files changed (1) hide show
  1. model/pipeline.py +13 -4
model/pipeline.py CHANGED
@@ -13,7 +13,7 @@ import torch
13
  import torch._dynamo
14
  from diffusers.pipelines.wan.pipeline_wan_i2v import WanImageToVideoPipeline
15
  from PIL import Image
16
- from torchao.quantization import Float8DynamicActivationFloat8WeightConfig, Int8WeightOnlyConfig, quantize_
17
 
18
  MODEL_ID = "thornmaze/WAMU_v3_WAN2.2_I2V_LIGHTNING"
19
 
@@ -41,12 +41,21 @@ def load_pipeline() -> WanImageToVideoPipeline:
41
  pipe = WanImageToVideoPipeline.from_pretrained(MODEL_ID, torch_dtype=torch.bfloat16).to("cuda")
42
 
43
  # Quantize to shrink the ~69GB bf16 pipe (drops per-call ZeroGPU init time and peak memory).
44
- # Independent of AOT compilation (FR-6/C-2, deferred) — same recipe as the reference project.
 
 
 
 
 
 
 
 
 
45
  quantize_(pipe.text_encoder, Int8WeightOnlyConfig())
46
  torch._dynamo.reset()
47
- quantize_(pipe.transformer, Float8DynamicActivationFloat8WeightConfig())
48
  torch._dynamo.reset()
49
- quantize_(pipe.transformer_2, Float8DynamicActivationFloat8WeightConfig())
50
  torch._dynamo.reset()
51
 
52
  return pipe
 
13
  import torch._dynamo
14
  from diffusers.pipelines.wan.pipeline_wan_i2v import WanImageToVideoPipeline
15
  from PIL import Image
16
+ from torchao.quantization import Float8WeightOnlyConfig, Int8WeightOnlyConfig, quantize_
17
 
18
  MODEL_ID = "thornmaze/WAMU_v3_WAN2.2_I2V_LIGHTNING"
19
 
 
41
  pipe = WanImageToVideoPipeline.from_pretrained(MODEL_ID, torch_dtype=torch.bfloat16).to("cuda")
42
 
43
  # Quantize to shrink the ~69GB bf16 pipe (drops per-call ZeroGPU init time and peak memory).
44
+ # Independent of AOT compilation (FR-6/C-2, deferred).
45
+ #
46
+ # Weight-only (not dynamic-activation) float8 for the transformers: the reference project uses
47
+ # Float8DynamicActivationFloat8WeightConfig, but only under spaces.aoti_load() (AOT-compiled),
48
+ # which pre-plans memory for the per-layer activation-quantize casts that config adds. Running
49
+ # that same config eagerly (as we do here, AOT deferred per #3) left every one of those casts as
50
+ # an unplanned CUDA allocation, which reliably crashed ZeroGPU with an NVML assert in the
51
+ # caching allocator (see #35). Weight-only quantization keeps the weight-memory savings without
52
+ # any runtime activation-quantization allocations, at the cost of the fp8x8 matmul speedup —
53
+ # revisit once AOT compilation (#36) makes the dynamic-activation config safe again.
54
  quantize_(pipe.text_encoder, Int8WeightOnlyConfig())
55
  torch._dynamo.reset()
56
+ quantize_(pipe.transformer, Float8WeightOnlyConfig())
57
  torch._dynamo.reset()
58
+ quantize_(pipe.transformer_2, Float8WeightOnlyConfig())
59
  torch._dynamo.reset()
60
 
61
  return pipe