Spaces:
Running on Zero
Sync develop into main (#34)
Browse files* Check Upscale 4x by default and drop model name from its label (#19)
Simplifies the option to a plain checkbox toggle without exposing the
underlying model name in the UI.
Claude-Session: https://claude.ai/code/session_01W9GLmk1R3mFfb7otDoGiWX
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* Add detailed export/traceback logging to diagnose post-pipeline UI errors (#20)
The Space logs cut off right after the pipeline stages finish, with no
trace of what happens next in generate_video() (export_to_video and the
run_inference call boundary), even though the UI shows a generic Error.
Log frame metadata and full tracebacks at those points so the next
failure is diagnosable from the Space logs alone.
* Pre-download RIFE/upscale weights at startup instead of inside @spaces.GPU (#22)
The RIFE and 4xLSDIRCompact weight fetches are pure network/disk I/O with no
CUDA dependency, but were only downloaded lazily on first use inside the
metered @spaces.GPU allocation — costing a few seconds of ZeroGPU quota on
every cold container start. Move both fetches to app startup (alongside
load_pipeline()), mirroring how the main video model is already preloaded.
Claude-Session: https://claude.ai/code/session_011keM1seSWsBZqPHrzKTeCG
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* Batch frames and same-shaped tiles through the upscale model (#23)
upscale_frames processed one frame at a time, and within each frame
_tile_process called the model once per tile (batch=1 every time). With
SRVGGNetCompact this leaves most of the GPU idle per call, since a single
~256px tile is far under typical GPU parallel capacity.
Group frames of equal size into batches (FRAME_BATCH_SIZE) before tiling,
and within _tile_process group tile jobs by input-patch shape (interior
tiles are all the same size; only the last row/column differs, clipped
against the image edge) so each model() call is one batched forward pass
over same-shaped patches, chunked to MAX_TILE_BATCH to bound memory. The
per-tile crop/placement math is unchanged — verified index-for-index
equivalent to the prior per-tile loop via a numpy simulation with a
synthetic nonlinear "model" across several batch/ragged-size cases.
No GPU available in this environment to benchmark; needs validation on
the dev Space before merge.
* Increase upscale tile size from 256 to 512 (#24)
Larger tiles mean fewer, bigger model() calls per frame, which cuts
pad/crop overhead and Python/kernel-launch count and improves GPU
occupancy for this small SRVGGNetCompact network. 512 is still
conservative for a 16GB+ Spaces GPU (T4/L4/A10G).
No GPU available in this environment to benchmark memory headroom or
speedup; needs validation on the dev Space before merge.
* torch.compile the upscale model on CUDA (#26)
Wraps the model with torch.compile(dynamic=True) on the CUDA path for
kernel fusion. dynamic=True avoids a fresh recompile for every distinct
tile shape (interior tiles vs. the clipped last row/column) — there's
only a handful of those per video, but without it each new shape would
trigger its own recompile. Wrapped in try/except: if torch.compile fails
to initialize (e.g. missing/incompatible triton on the target GPU image),
the eager model is used as-is.
No GPU available in this environment to benchmark, and this is the
riskiest of the four upscale perf changes (compile warmup cost vs.
steady-state speedup, triton availability on the Space's GPU image) —
needs careful validation on the dev Space before merge.
* Log when torch.compile falls back to the eager upscale model (#31)
The try/except guard added in #26 silently swallowed torch.compile
initialization failures (e.g. missing/incompatible triton on the Space's
GPU image), so a fallback to the eager model would be invisible in the
logs — leaving no way to tell, from the deployed Space's console output,
whether the compiled path is actually active. Print the exception instead
of passing silently.
* Add liability disclaimer to the Space UI (#30)
Notes that outputs are provided "as is" and that users are solely
responsible for generated content and how it's used/shared.
Claude-Session: https://claude.ai/code/session_015woCw9matvyXW3Xnx2JwXG
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* Stop Inductor compilation progress from leaking into the Gradio UI (#33)
gr.Progress(track_tqdm=True) monkey-patches tqdm globally, so it was
picking up TorchInductor's own "Inductor Compilation" progress bars
(triggered by torch.compile on the upscale model, and by the quantized
transformer/text_encoder) regardless of their own disable flag. Drop
track_tqdm and report the generation stage's progress explicitly via
callback_on_step_end instead, matching the interpolation/upscale stages.
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
|
@@ -161,11 +161,15 @@ def run_inference(
|
|
| 161 |
seed,
|
| 162 |
frame_multiplier,
|
| 163 |
upscale_output,
|
| 164 |
-
progress=gr.Progress(
|
| 165 |
):
|
| 166 |
print_infer_start(prompt, negative_prompt, seed, steps, guidance_scale, frame_multiplier, upscale_output)
|
| 167 |
t_start = _time.perf_counter()
|
| 168 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 169 |
print_stage_start("generation")
|
| 170 |
t0 = _time.perf_counter()
|
| 171 |
try:
|
|
@@ -180,6 +184,7 @@ def run_inference(
|
|
| 180 |
num_inference_steps=int(steps),
|
| 181 |
generator=torch.Generator(device="cuda").manual_seed(seed),
|
| 182 |
output_type="np",
|
|
|
|
| 183 |
)
|
| 184 |
except Exception as e:
|
| 185 |
print_stage_error("generation", e)
|
|
@@ -246,7 +251,7 @@ def generate_video(
|
|
| 246 |
randomize_seed,
|
| 247 |
frame_multiplier,
|
| 248 |
upscale_output,
|
| 249 |
-
progress=gr.Progress(
|
| 250 |
):
|
| 251 |
if input_image is None:
|
| 252 |
raise gr.Error("Please upload an input image.")
|
|
|
|
| 161 |
seed,
|
| 162 |
frame_multiplier,
|
| 163 |
upscale_output,
|
| 164 |
+
progress=gr.Progress(),
|
| 165 |
):
|
| 166 |
print_infer_start(prompt, negative_prompt, seed, steps, guidance_scale, frame_multiplier, upscale_output)
|
| 167 |
t_start = _time.perf_counter()
|
| 168 |
|
| 169 |
+
def _report_generation_progress(pipe_, step_index, timestep, callback_kwargs):
|
| 170 |
+
progress(0.7 * (step_index + 1) / int(steps), desc=f"Generating ({step_index + 1}/{int(steps)} steps)...")
|
| 171 |
+
return callback_kwargs
|
| 172 |
+
|
| 173 |
print_stage_start("generation")
|
| 174 |
t0 = _time.perf_counter()
|
| 175 |
try:
|
|
|
|
| 184 |
num_inference_steps=int(steps),
|
| 185 |
generator=torch.Generator(device="cuda").manual_seed(seed),
|
| 186 |
output_type="np",
|
| 187 |
+
callback_on_step_end=_report_generation_progress,
|
| 188 |
)
|
| 189 |
except Exception as e:
|
| 190 |
print_stage_error("generation", e)
|
|
|
|
| 251 |
randomize_seed,
|
| 252 |
frame_multiplier,
|
| 253 |
upscale_output,
|
| 254 |
+
progress=gr.Progress(),
|
| 255 |
):
|
| 256 |
if input_image is None:
|
| 257 |
raise gr.Error("Please upload an input image.")
|