Text-to-Image
Diffusers
Safetensors
QwenImage21Pipeline
qwen-image
qwen-image-2.1
nvfp4
svdquant
nunchaku
blackwell
image-editing
8-bit precision
Instructions to use joseplcam/Qwen-Image-2.1-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use joseplcam/Qwen-Image-2.1-NVFP4 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("joseplcam/Qwen-Image-2.1-NVFP4", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Card: eager is the better default; torch.compile recompiles on mixed workloads
Browse files
README.md
CHANGED
|
@@ -39,7 +39,6 @@ from diffusers import DiffusionPipeline
|
|
| 39 |
pipe = DiffusionPipeline.from_pretrained(
|
| 40 |
"joseplcam/Qwen-Image-2.1-NVFP4", dtype=torch.bfloat16, trust_remote_code=True
|
| 41 |
).to("cuda")
|
| 42 |
-
pipe.transformer.compile_repeated_blocks() # optional: ~1.2x faster; the first call compiles (~10 s)
|
| 43 |
|
| 44 |
image = pipe("A capybara reading a book by candlelight", height=512, width=512, num_inference_steps=40).images[0]
|
| 45 |
edited = pipe("Make it night time with moonlight", image=image, output_resolution=512, num_inference_steps=40).images[0]
|
|
@@ -117,16 +116,20 @@ Peak VRAM is PyTorch's peak allocation.
|
|
| 117 |
|
| 118 |
Batching several prompts gives no extra throughput in Diffusers on this GPU; it is saturated at batch 1.
|
| 119 |
|
| 120 |
-
Runtime options (text-to-image, 512x512, 40 steps, mean of 3 runs after warmup):
|
| 121 |
|
| 122 |
| Setup | BF16 (official) | This repo |
|
| 123 |
|---|---|---|
|
| 124 |
| Eager, PyTorch SDPA | 3.20 s | 1.79 s |
|
| 125 |
| Eager, cuDNN SDPA | 3.26 s | 1.88 s |
|
| 126 |
-
|
|
| 127 |
-
|
| 128 |
-
-
|
| 129 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 130 |
- CUDA graphs (`mode="reduce-overhead"`) do not work: the prefix KV cache keeps tensors that graph replays
|
| 131 |
overwrite, and the Nunchaku kernels cannot be captured (`cudaErrorStreamCaptureInvalidated`). With the KV
|
| 132 |
cache off, BF16 gained only 6% from them.
|
|
|
|
| 39 |
pipe = DiffusionPipeline.from_pretrained(
|
| 40 |
"joseplcam/Qwen-Image-2.1-NVFP4", dtype=torch.bfloat16, trust_remote_code=True
|
| 41 |
).to("cuda")
|
|
|
|
| 42 |
|
| 43 |
image = pipe("A capybara reading a book by candlelight", height=512, width=512, num_inference_steps=40).images[0]
|
| 44 |
edited = pipe("Make it night time with moonlight", image=image, output_resolution=512, num_inference_steps=40).images[0]
|
|
|
|
| 116 |
|
| 117 |
Batching several prompts gives no extra throughput in Diffusers on this GPU; it is saturated at batch 1.
|
| 118 |
|
| 119 |
+
Runtime options (text-to-image, 512x512, 40 steps, one fixed prompt, mean of 3 runs after warmup):
|
| 120 |
|
| 121 |
| Setup | BF16 (official) | This repo |
|
| 122 |
|---|---|---|
|
| 123 |
| Eager, PyTorch SDPA | 3.20 s | 1.79 s |
|
| 124 |
| Eager, cuDNN SDPA | 3.26 s | 1.88 s |
|
| 125 |
+
| `compile_repeated_blocks()`, PyTorch SDPA | 2.99 s | 1.50 s |
|
| 126 |
+
|
| 127 |
+
- **Use eager (the default) for mixed workloads.** `torch.compile` only wins while every call has the same
|
| 128 |
+
shapes. Prompt length and the reference image's aspect ratio change the prefix KV cache, so varied
|
| 129 |
+
text-to-image prompts and edits keep recompiling until TorchDynamo's recompile limit, after which new shapes
|
| 130 |
+
run uncompiled anyway. In a 12-call mix (6 prompts of different lengths, 6 edits with different reference
|
| 131 |
+
shapes) eager took 37.2 s in total, `compile_repeated_blocks()` 44.3 s and `dynamic=True` 56.5 s; steady-state
|
| 132 |
+
edits were 2.0 s, 2.0 s and 1.7 s. Compile is only worth it for fixed-shape batch jobs.
|
| 133 |
- CUDA graphs (`mode="reduce-overhead"`) do not work: the prefix KV cache keeps tensors that graph replays
|
| 134 |
overwrite, and the Nunchaku kernels cannot be captured (`cudaErrorStreamCaptureInvalidated`). With the KV
|
| 135 |
cache off, BF16 gained only 6% from them.
|