joseplcam commited on
Commit
51aa90f
·
verified ·
1 Parent(s): 4cc13ed

Card: eager is the better default; torch.compile recompiles on mixed workloads

Browse files
Files changed (1) hide show
  1. README.md +9 -6
README.md CHANGED
@@ -39,7 +39,6 @@ from diffusers import DiffusionPipeline
39
  pipe = DiffusionPipeline.from_pretrained(
40
  "joseplcam/Qwen-Image-2.1-NVFP4", dtype=torch.bfloat16, trust_remote_code=True
41
  ).to("cuda")
42
- pipe.transformer.compile_repeated_blocks() # optional: ~1.2x faster; the first call compiles (~10 s)
43
 
44
  image = pipe("A capybara reading a book by candlelight", height=512, width=512, num_inference_steps=40).images[0]
45
  edited = pipe("Make it night time with moonlight", image=image, output_resolution=512, num_inference_steps=40).images[0]
@@ -117,16 +116,20 @@ Peak VRAM is PyTorch's peak allocation.
117
 
118
  Batching several prompts gives no extra throughput in Diffusers on this GPU; it is saturated at batch 1.
119
 
120
- Runtime options (text-to-image, 512x512, 40 steps, mean of 3 runs after warmup):
121
 
122
  | Setup | BF16 (official) | This repo |
123
  |---|---|---|
124
  | Eager, PyTorch SDPA | 3.20 s | 1.79 s |
125
  | Eager, cuDNN SDPA | 3.26 s | 1.88 s |
126
- | **`compile_repeated_blocks()`, PyTorch SDPA** | 2.99 s | **1.50 s** (2.1x vs BF16 eager) |
127
-
128
- - `torch.compile` without CUDA graphs is the fastest setup here and does not change the images beyond the
129
- run-to-run variation described below.
 
 
 
 
130
  - CUDA graphs (`mode="reduce-overhead"`) do not work: the prefix KV cache keeps tensors that graph replays
131
  overwrite, and the Nunchaku kernels cannot be captured (`cudaErrorStreamCaptureInvalidated`). With the KV
132
  cache off, BF16 gained only 6% from them.
 
39
  pipe = DiffusionPipeline.from_pretrained(
40
  "joseplcam/Qwen-Image-2.1-NVFP4", dtype=torch.bfloat16, trust_remote_code=True
41
  ).to("cuda")
 
42
 
43
  image = pipe("A capybara reading a book by candlelight", height=512, width=512, num_inference_steps=40).images[0]
44
  edited = pipe("Make it night time with moonlight", image=image, output_resolution=512, num_inference_steps=40).images[0]
 
116
 
117
  Batching several prompts gives no extra throughput in Diffusers on this GPU; it is saturated at batch 1.
118
 
119
+ Runtime options (text-to-image, 512x512, 40 steps, one fixed prompt, mean of 3 runs after warmup):
120
 
121
  | Setup | BF16 (official) | This repo |
122
  |---|---|---|
123
  | Eager, PyTorch SDPA | 3.20 s | 1.79 s |
124
  | Eager, cuDNN SDPA | 3.26 s | 1.88 s |
125
+ | `compile_repeated_blocks()`, PyTorch SDPA | 2.99 s | 1.50 s |
126
+
127
+ - **Use eager (the default) for mixed workloads.** `torch.compile` only wins while every call has the same
128
+ shapes. Prompt length and the reference image's aspect ratio change the prefix KV cache, so varied
129
+ text-to-image prompts and edits keep recompiling until TorchDynamo's recompile limit, after which new shapes
130
+ run uncompiled anyway. In a 12-call mix (6 prompts of different lengths, 6 edits with different reference
131
+ shapes) eager took 37.2 s in total, `compile_repeated_blocks()` 44.3 s and `dynamic=True` 56.5 s; steady-state
132
+ edits were 2.0 s, 2.0 s and 1.7 s. Compile is only worth it for fixed-shape batch jobs.
133
  - CUDA graphs (`mode="reduce-overhead"`) do not work: the prefix KV cache keeps tensors that graph replays
134
  overwrite, and the Nunchaku kernels cannot be captured (`cudaErrorStreamCaptureInvalidated`). With the KV
135
  cache off, BF16 gained only 6% from them.