someone-in-the-world Claude Sonnet 5 commited on
Commit
5ab12ca
·
unverified ·
1 Parent(s): 124fb40

Sync develop into main (#34)

Browse files

* Check Upscale 4x by default and drop model name from its label (#19)

Simplifies the option to a plain checkbox toggle without exposing the
underlying model name in the UI.


Claude-Session: https://claude.ai/code/session_01W9GLmk1R3mFfb7otDoGiWX

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* Add detailed export/traceback logging to diagnose post-pipeline UI errors (#20)

The Space logs cut off right after the pipeline stages finish, with no
trace of what happens next in generate_video() (export_to_video and the
run_inference call boundary), even though the UI shows a generic Error.
Log frame metadata and full tracebacks at those points so the next
failure is diagnosable from the Space logs alone.

* Pre-download RIFE/upscale weights at startup instead of inside @spaces.GPU (#22)

The RIFE and 4xLSDIRCompact weight fetches are pure network/disk I/O with no
CUDA dependency, but were only downloaded lazily on first use inside the
metered @spaces.GPU allocation — costing a few seconds of ZeroGPU quota on
every cold container start. Move both fetches to app startup (alongside
load_pipeline()), mirroring how the main video model is already preloaded.


Claude-Session: https://claude.ai/code/session_011keM1seSWsBZqPHrzKTeCG

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* Batch frames and same-shaped tiles through the upscale model (#23)

upscale_frames processed one frame at a time, and within each frame
_tile_process called the model once per tile (batch=1 every time). With
SRVGGNetCompact this leaves most of the GPU idle per call, since a single
~256px tile is far under typical GPU parallel capacity.

Group frames of equal size into batches (FRAME_BATCH_SIZE) before tiling,
and within _tile_process group tile jobs by input-patch shape (interior
tiles are all the same size; only the last row/column differs, clipped
against the image edge) so each model() call is one batched forward pass
over same-shaped patches, chunked to MAX_TILE_BATCH to bound memory. The
per-tile crop/placement math is unchanged — verified index-for-index
equivalent to the prior per-tile loop via a numpy simulation with a
synthetic nonlinear "model" across several batch/ragged-size cases.

No GPU available in this environment to benchmark; needs validation on
the dev Space before merge.

* Increase upscale tile size from 256 to 512 (#24)

Larger tiles mean fewer, bigger model() calls per frame, which cuts
pad/crop overhead and Python/kernel-launch count and improves GPU
occupancy for this small SRVGGNetCompact network. 512 is still
conservative for a 16GB+ Spaces GPU (T4/L4/A10G).

No GPU available in this environment to benchmark memory headroom or
speedup; needs validation on the dev Space before merge.

* torch.compile the upscale model on CUDA (#26)

Wraps the model with torch.compile(dynamic=True) on the CUDA path for
kernel fusion. dynamic=True avoids a fresh recompile for every distinct
tile shape (interior tiles vs. the clipped last row/column) — there's
only a handful of those per video, but without it each new shape would
trigger its own recompile. Wrapped in try/except: if torch.compile fails
to initialize (e.g. missing/incompatible triton on the target GPU image),
the eager model is used as-is.

No GPU available in this environment to benchmark, and this is the
riskiest of the four upscale perf changes (compile warmup cost vs.
steady-state speedup, triton availability on the Space's GPU image) —
needs careful validation on the dev Space before merge.

* Log when torch.compile falls back to the eager upscale model (#31)

The try/except guard added in #26 silently swallowed torch.compile
initialization failures (e.g. missing/incompatible triton on the Space's
GPU image), so a fallback to the eager model would be invisible in the
logs — leaving no way to tell, from the deployed Space's console output,
whether the compiled path is actually active. Print the exception instead
of passing silently.

* Add liability disclaimer to the Space UI (#30)

Notes that outputs are provided "as is" and that users are solely
responsible for generated content and how it's used/shared.


Claude-Session: https://claude.ai/code/session_015woCw9matvyXW3Xnx2JwXG

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* Stop Inductor compilation progress from leaking into the Gradio UI (#33)

gr.Progress(track_tqdm=True) monkey-patches tqdm globally, so it was
picking up TorchInductor's own "Inductor Compilation" progress bars
(triggered by torch.compile on the upscale model, and by the quantized
transformer/text_encoder) regardless of their own disable flag. Drop
track_tqdm and report the generation stage's progress explicitly via
callback_on_step_end instead, matching the interpolation/upscale stages.

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

Files changed (1) hide show
  1. app.py +7 -2
app.py CHANGED
@@ -161,11 +161,15 @@ def run_inference(
161
  seed,
162
  frame_multiplier,
163
  upscale_output,
164
- progress=gr.Progress(track_tqdm=True),
165
  ):
166
  print_infer_start(prompt, negative_prompt, seed, steps, guidance_scale, frame_multiplier, upscale_output)
167
  t_start = _time.perf_counter()
168
 
 
 
 
 
169
  print_stage_start("generation")
170
  t0 = _time.perf_counter()
171
  try:
@@ -180,6 +184,7 @@ def run_inference(
180
  num_inference_steps=int(steps),
181
  generator=torch.Generator(device="cuda").manual_seed(seed),
182
  output_type="np",
 
183
  )
184
  except Exception as e:
185
  print_stage_error("generation", e)
@@ -246,7 +251,7 @@ def generate_video(
246
  randomize_seed,
247
  frame_multiplier,
248
  upscale_output,
249
- progress=gr.Progress(track_tqdm=True),
250
  ):
251
  if input_image is None:
252
  raise gr.Error("Please upload an input image.")
 
161
  seed,
162
  frame_multiplier,
163
  upscale_output,
164
+ progress=gr.Progress(),
165
  ):
166
  print_infer_start(prompt, negative_prompt, seed, steps, guidance_scale, frame_multiplier, upscale_output)
167
  t_start = _time.perf_counter()
168
 
169
+ def _report_generation_progress(pipe_, step_index, timestep, callback_kwargs):
170
+ progress(0.7 * (step_index + 1) / int(steps), desc=f"Generating ({step_index + 1}/{int(steps)} steps)...")
171
+ return callback_kwargs
172
+
173
  print_stage_start("generation")
174
  t0 = _time.perf_counter()
175
  try:
 
184
  num_inference_steps=int(steps),
185
  generator=torch.Generator(device="cuda").manual_seed(seed),
186
  output_type="np",
187
+ callback_on_step_end=_report_generation_progress,
188
  )
189
  except Exception as e:
190
  print_stage_error("generation", e)
 
251
  randomize_seed,
252
  frame_multiplier,
253
  upscale_output,
254
+ progress=gr.Progress(),
255
  ):
256
  if input_image is None:
257
  raise gr.Error("Please upload an input image.")