Text-to-Image
Diffusers
Safetensors
QwenImage21Pipeline
qwen-image
qwen-image-2.1
nvfp4
svdquant
nunchaku
blackwell
image-editing
8-bit precision
Instructions to use joseplcam/Qwen-Image-2.1-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use joseplcam/Qwen-Image-2.1-NVFP4 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("joseplcam/Qwen-Image-2.1-NVFP4", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Replace with SVDQuant+GPTQ NVFP4 DiT and text encoder, native in Diffusers
Browse files- NOTICE +25 -15
- README.md +123 -102
- model_index.json +5 -5
- text_encoder/config.json +95 -0
- text_encoder/model.safetensors +2 -2
- text_encoder/modeling_nunchaku_qwen3vl.py +47 -0
- text_encoder/nunchaku_kernels.py +19 -0
- tools/evaluate.py +103 -0
- tools/modeling_nunchaku_qwen3vl.py +47 -0
- tools/modeling_nunchaku_qwenimage21.py +14 -0
- tools/nunchaku_kernels.py +19 -0
- tools/package.py +81 -0
- tools/parrot_test.py +69 -0
- tools/quantize_dit.py +144 -0
- tools/quantize_text_encoder.py +121 -0
- transformer/config.json +207 -41
- transformer/diffusion_pytorch_model.safetensors +2 -2
- transformer/modeling_nunchaku_qwenimage21.py +14 -0
- transformer/nunchaku_kernels.py +19 -0
- vae/config.json +1 -1
- vae/diffusion_pytorch_model.safetensors +2 -2
NOTICE
CHANGED
|
@@ -2,22 +2,32 @@ Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 H
|
|
| 2 |
|
| 3 |
Built with Qwen.
|
| 4 |
|
| 5 |
-
This repository redistributes, for non-commercial research and evaluation use only, files derived from
|
|
|
|
| 6 |
|
| 7 |
-
|
| 8 |
-
model_index.json, LICENSE, processor/, scheduler/, vae/, text_encoder/config.json,
|
| 9 |
-
text_encoder/generation_config.json. Unmodified.
|
| 10 |
|
| 11 |
-
|
| 12 |
-
transformer/config.json, transformer/diffusion_pytorch_model.safetensors. Unmodified
|
| 13 |
-
(sha256 a6c3080b764344729ef95846faa3a0a01eb36bdef93f785ed07f4ac5751ad76a).
|
| 14 |
|
| 15 |
-
-
|
| 16 |
-
|
| 17 |
-
|
|
|
|
|
|
|
| 18 |
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
|
| 3 |
Built with Qwen.
|
| 4 |
|
| 5 |
+
This repository redistributes, for non-commercial research and evaluation use only, files derived from
|
| 6 |
+
Qwen/Qwen-Image-2.1 (revision 790c92633540aa0cb11d9abf19eb46d861714758).
|
| 7 |
|
| 8 |
+
Unmodified files: LICENSE, processor/, scheduler/, text_encoder/generation_config.json.
|
|
|
|
|
|
|
| 9 |
|
| 10 |
+
MODIFIED FILES (changed by joseplcam):
|
|
|
|
|
|
|
| 11 |
|
| 12 |
+
- transformer/diffusion_pytorch_model.safetensors, transformer/config.json
|
| 13 |
+
Quantized from the official BF16 transformer with SVDQuant + GPTQ (NVFP4 weights and activations,
|
| 14 |
+
rank-32 BF16 low-rank branch) using diffuse-compressor. Blocks 0, 1, 30, 31 and the global
|
| 15 |
+
modulation are the official BF16 tensors. config.json gained a `quantization_config` entry.
|
| 16 |
+
sha256 0304fc3e84b13663b439601380066754ef9b92025e159275abe163bc5f1ceeb4
|
| 17 |
|
| 18 |
+
- text_encoder/model.safetensors, text_encoder/config.json
|
| 19 |
+
Quantized from the official BF16 Qwen3-VL text encoder with the same method: the MLP projections
|
| 20 |
+
(gate/up/down) of decoder layers 4-31, rank-128 low-rank branch. All attention projections, decoder
|
| 21 |
+
layers 0-3 and 32-35, the vision tower, embeddings and lm_head are the official BF16 tensors.
|
| 22 |
+
config.json gained a `nunchaku_lite` entry.
|
| 23 |
+
sha256 83c4462b5ce6b98e2708ca53031e2f9a60a8e402c4c4a0ccd19cb2df8e817ae2
|
| 24 |
+
|
| 25 |
+
- vae/diffusion_pytorch_model.safetensors, vae/config.json
|
| 26 |
+
The official VAE weights cast from FP32 to BF16.
|
| 27 |
+
sha256 71879ffd5321e6d10c3c87513e2b474b1252efa7f3dec2969214a9bf06a6dd5c
|
| 28 |
+
|
| 29 |
+
- model_index.json
|
| 30 |
+
The `text_encoder` and `transformer` entries point to the custom classes below.
|
| 31 |
+
|
| 32 |
+
NEW FILES: text_encoder/modeling_nunchaku_qwen3vl.py, transformer/modeling_nunchaku_qwenimage21.py,
|
| 33 |
+
text_encoder/nunchaku_kernels.py, transformer/nunchaku_kernels.py.
|
README.md
CHANGED
|
@@ -10,129 +10,150 @@ tags:
|
|
| 10 |
- qwen-image
|
| 11 |
- qwen-image-2.1
|
| 12 |
- nvfp4
|
| 13 |
-
-
|
|
|
|
| 14 |
- blackwell
|
| 15 |
- text-to-image
|
| 16 |
- image-editing
|
| 17 |
---
|
| 18 |
|
| 19 |
-
# Qwen-Image-2.1-NVFP4 (DiT + text encoder
|
| 20 |
|
| 21 |
**Built with Qwen.** Non-commercial research and evaluation use only (Qwen Research License, see `LICENSE` and `NOTICE`).
|
| 22 |
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
combination, the small key-rename that makes the encoder loadable by SGLang, and the measurements below.
|
| 28 |
|
| 29 |
-
|
|
|
|
|
|
|
| 30 |
|
| 31 |
-
##
|
| 32 |
-
|
| 33 |
-
| Path | Precision | Source | Changes |
|
| 34 |
-
|---|---|---|---|
|
| 35 |
-
| `transformer/` (7B single-stream DiT) | NVFP4 W4A4, group 16, FP8 block scales. Blocks 0, 1, 30, 31 and the embeddings, modulation and output layers stay BF16 | [HangGlidersRule/Darkstar-Qwen-Image-2.1-Base-ModelOpt-W4A4-NVFP4](https://huggingface.co/HangGlidersRule/Darkstar-Qwen-Image-2.1-Base-ModelOpt-W4A4-NVFP4) @ `3b7aa84`. NVIDIA ModelOpt export of the base model | None (byte-identical) |
|
| 36 |
-
| `text_encoder/model.safetensors` (Qwen3-VL 8B) | NVFP4 for the 252 language-model linear layers; vision tower, embeddings and `lm_head` stay BF16 | [BennyDaBall/Qwen-Image-2.1-NVFP4](https://huggingface.co/BennyDaBall/Qwen-Image-2.1-NVFP4) @ `1a38d44`, file `text_encoders/qwen3vl_8b_nvfp4.safetensors`. ComfyUI-style NVFP4 converted from the official BF16 weights | Tensor names renamed to the Hugging Face Qwen3-VL layout (`model.layers.*` → `model.language_model.layers.*`, same for `embed_tokens` and `norm`). Data unchanged |
|
| 37 |
-
| `vae/`, `processor/`, `scheduler/`, `model_index.json`, `text_encoder/*.json`, `LICENSE` | BF16 / FP32 as released | [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) @ `790c926` | None |
|
| 38 |
-
|
| 39 |
-
Why these two parts:
|
| 40 |
-
|
| 41 |
-
- The DiT is a ModelOpt NVFP4 export. SGLang detects it from `transformer/config.json` and runs it with FlashInfer's
|
| 42 |
-
CUTLASS FP4 GEMM, which has an SM120 path (RTX 50 series, RTX PRO 6000). Keeping the first two and last two blocks in
|
| 43 |
-
BF16 preserves quality: its author reports parity with BF16 on a 25-case evaluation.
|
| 44 |
-
- The encoder is the only published NVFP4 Qwen3-VL 8B built from the official weights (not an abliterated variant).
|
| 45 |
-
SGLang runs its Comfy NVFP4 layers natively via `comfy-kitchen` FP4 matmuls. The rename was needed because SGLang's
|
| 46 |
-
native Qwen3-VL encoder expects the Hugging Face tensor names.
|
| 47 |
-
|
| 48 |
-
Total size is about 14 GB, against about 32 GB for the BF16 release.
|
| 49 |
-
|
| 50 |
-
## Usage (SGLang Diffusion)
|
| 51 |
-
|
| 52 |
-
Tested with SGLang commit `ffac53d779c08dcdab2d07e5e2a41dba83f0e65c`, PyTorch 2.13.0+cu130, FlashInfer 0.6.18 and
|
| 53 |
-
`comfy-kitchen==0.2.35`. You need a Blackwell GPU (compute capability 10.0 or 12.0).
|
| 54 |
-
|
| 55 |
-
SGLang only recognizes the Comfy NVFP4 encoder when its weights path points at the file, so pass it explicitly:
|
| 56 |
|
| 57 |
```python
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
repo = "joseplcam/Qwen-Image-2.1-NVFP4"
|
| 61 |
-
with DiffGenerator.from_pretrained(
|
| 62 |
-
model_path=repo,
|
| 63 |
-
component_weights_paths={"text_encoder": f"{repo}/text_encoder/model.safetensors"},
|
| 64 |
-
performance_mode="speed", # keep every component resident on the GPU
|
| 65 |
-
attention_backend="torch_sdpa", # "sage_attn" is faster at 1024x1024 and above
|
| 66 |
-
) as generator:
|
| 67 |
-
result = generator.generate(sampling_params_kwargs=dict(
|
| 68 |
-
prompt="A capybara reading a book by candlelight",
|
| 69 |
-
width=512, height=512, num_inference_steps=40, guidance_scale=1.0, seed=42,
|
| 70 |
-
))
|
| 71 |
-
```
|
| 72 |
-
|
| 73 |
-
Add `image_path="input.png"` for editing. The same arguments work with a local clone of this repository.
|
| 74 |
|
| 75 |
-
|
|
|
|
|
|
|
| 76 |
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
--component-weights-paths.text_encoder joseplcam/Qwen-Image-2.1-NVFP4/text_encoder/model.safetensors \
|
| 80 |
-
--performance-mode speed --batching-max-size 8 --batching-delay-ms 50
|
| 81 |
```
|
| 82 |
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
-
|
| 86 |
-
|
| 87 |
-
-
|
| 88 |
-
|
| 89 |
-
-
|
| 90 |
-
`
|
| 91 |
-
|
| 92 |
-
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
|
| 96 |
-
|
| 97 |
-
|
| 98 |
-
|
| 99 |
-
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
|
| 103 |
-
|
| 104 |
-
|
| 105 |
-
|
| 106 |
-
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
|
|
| 113 |
-
|
| 114 |
-
|
| 115 |
-
|
| 116 |
-
|
| 117 |
-
|
| 118 |
-
|
| 119 |
-
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 123 |
|
| 124 |
## Limitations
|
| 125 |
|
|
|
|
|
|
|
|
|
|
| 126 |
- Non-commercial research and evaluation use only.
|
| 127 |
-
|
| 128 |
-
|
| 129 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 130 |
|
| 131 |
## Credits and license
|
| 132 |
|
| 133 |
-
- Qwen team: Qwen-Image
|
| 134 |
-
-
|
| 135 |
-
|
|
|
|
| 136 |
|
| 137 |
-
Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory
|
| 138 |
-
Co., Ltd. All Rights Reserved. See `LICENSE` and `NOTICE`, which
|
|
|
|
| 10 |
- qwen-image
|
| 11 |
- qwen-image-2.1
|
| 12 |
- nvfp4
|
| 13 |
+
- svdquant
|
| 14 |
+
- nunchaku
|
| 15 |
- blackwell
|
| 16 |
- text-to-image
|
| 17 |
- image-editing
|
| 18 |
---
|
| 19 |
|
| 20 |
+
# Qwen-Image-2.1-NVFP4 (SVDQuant, DiT + text encoder, native in Diffusers)
|
| 21 |
|
| 22 |
**Built with Qwen.** Non-commercial research and evaluation use only (Qwen Research License, see `LICENSE` and `NOTICE`).
|
| 23 |
|
| 24 |
+
An NVFP4 quantization of [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1), made
|
| 25 |
+
from the official BF16 weights with SVDQuant + GPTQ and calibrated on both text-to-image prompts and
|
| 26 |
+
image edits. Both large components run on native Blackwell FP4 tensor cores, and the pipeline loads with a
|
| 27 |
+
single `DiffusionPipeline.from_pretrained` call.
|
|
|
|
| 28 |
|
| 29 |
+
At 512x512 on an RTX PRO 6000 Blackwell it is about **1.7x faster than BF16 and needs about 40% less VRAM**,
|
| 30 |
+
and it matches BF16 closely on text-to-image, typography and edits, including a hard recolor edit that
|
| 31 |
+
earlier 4-bit text encoders failed.
|
| 32 |
|
| 33 |
+
## Usage
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 34 |
|
| 35 |
```python
|
| 36 |
+
import torch
|
| 37 |
+
from diffusers import DiffusionPipeline
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
|
| 39 |
+
pipe = DiffusionPipeline.from_pretrained(
|
| 40 |
+
"joseplcam/Qwen-Image-2.1-NVFP4", dtype=torch.bfloat16, trust_remote_code=True
|
| 41 |
+
).to("cuda")
|
| 42 |
|
| 43 |
+
image = pipe("A capybara reading a book by candlelight", height=512, width=512, num_inference_steps=40).images[0]
|
| 44 |
+
edited = pipe("Make it night time with moonlight", image=image, output_resolution=512, num_inference_steps=40).images[0]
|
|
|
|
|
|
|
| 45 |
```
|
| 46 |
|
| 47 |
+
`trust_remote_code=True` loads two small files from this repository:
|
| 48 |
+
|
| 49 |
+
- `text_encoder/modeling_nunchaku_qwen3vl.py`: a `Qwen3VLForConditionalGeneration` subclass that swaps the
|
| 50 |
+
quantized linears for Diffusers' own `SVDQW4A4Linear` before loading the weights.
|
| 51 |
+
- `nunchaku_kernels.py` (in `text_encoder/` and `transformer/`): Diffusers loads the NVFP4 kernels from
|
| 52 |
+
`rootonchair/nunchaku-lite-kernels`, which is no longer downloadable. This file points that name at
|
| 53 |
+
[joseplcam/nunchaku-lite-kernels](https://huggingface.co/joseplcam/nunchaku-lite-kernels), an unmodified
|
| 54 |
+
build of the same open-source kernels, and sets `DIFFUSERS_TRUST_REMOTE_KERNELS=true` unless you already set
|
| 55 |
+
it. Set `LOCAL_KERNELS` yourself to use a different build.
|
| 56 |
+
|
| 57 |
+
Use a different seed for an edit than the one that generated its input image. Qwen-Image-2.1 returns an
|
| 58 |
+
over-sharpened copy that ignores the prompt when the edit starts from the same noise
|
| 59 |
+
([diffusers #14824](https://github.com/huggingface/diffusers/issues/14824)); this is a base-model behaviour.
|
| 60 |
+
|
| 61 |
+
### Requirements
|
| 62 |
+
|
| 63 |
+
- NVIDIA Blackwell GPU with compute capability 12.0 (RTX 50 series, RTX PRO 6000). The published kernel build
|
| 64 |
+
targets `sm_120a` only.
|
| 65 |
+
- Linux x86_64, Python 3.12, PyTorch 2.13 with CUDA 13.0, which the kernel build targets.
|
| 66 |
+
- Diffusers from `main` with Qwen-Image-2.1 and Nunchaku Lite support (tested at commit `0377f0c`),
|
| 67 |
+
`transformers>=5.12`, `accelerate`, `kernels>=0.14`.
|
| 68 |
+
|
| 69 |
+
## What is inside
|
| 70 |
+
|
| 71 |
+
| Component | Precision | Details |
|
| 72 |
+
|---|---|---|
|
| 73 |
+
| `transformer/` (7B DiT) | NVFP4 W4A4, group 16, FP8 block scales + BF16 rank-32 low-rank branch | All attention and MLP projections of blocks 2-29 (196 layers). Blocks 0, 1, 30, 31 and the global modulation stay BF16. |
|
| 74 |
+
| `text_encoder/` (Qwen3-VL 8B) | NVFP4 W4A4 + BF16 rank-128 low-rank branch | MLP projections (gate/up/down) of decoder layers 4-31 (84 layers). All attention projections, layers 0-3 and 32-35, the vision tower, embeddings and `lm_head` stay BF16. |
|
| 75 |
+
| `vae/` | BF16 | Official weights cast from FP32. |
|
| 76 |
+
| `processor/`, `scheduler/` | as released | Unmodified. |
|
| 77 |
+
|
| 78 |
+
Total download: about 17 GB, against about 32 GB for the BF16 release.
|
| 79 |
+
|
| 80 |
+
## How it was made
|
| 81 |
+
|
| 82 |
+
- **Method:** SVDQuant with GPTQ residual rounding, via [diffuse-compressor](https://github.com/rootonchair/diffuse-compressor)
|
| 83 |
+
(commit `0965874`). A low-rank BF16 branch absorbs the outliers of each weight and SmoothQuant-style
|
| 84 |
+
scaling migrates activation outliers; GPTQ then rounds the 4-bit residual using calibration statistics.
|
| 85 |
+
Activation scales are dynamic, so nothing is fixed to the calibration inputs.
|
| 86 |
+
- **Calibration data (128 samples per component):** 64 prompts from the qdiff prompt set and 64 image edits
|
| 87 |
+
from the train split of [VyoJ/NHR-Edit-Change_Only](https://huggingface.co/datasets/VyoJ/NHR-Edit-Change_Only).
|
| 88 |
+
Three of every four samples at 512x512, the rest at 1024x1024. The DiT saw 20 denoising steps per sample,
|
| 89 |
+
with the prefix KV cache disabled so the prompt and reference-image tokens pass through every step.
|
| 90 |
+
- **Sensitive layers:** the first and last two DiT blocks and the modulation stay BF16. For the text encoder,
|
| 91 |
+
quantizing attention hurt edits most. With every encoder linear quantized, the "turn the parrots blue"
|
| 92 |
+
recolor below worked in 1 of 8 seeds (rank 32) or 4-5 of 8 (rank 128); quantizing only the MLPs with rank
|
| 93 |
+
128 brought it to 7 of 8, the same as BF16.
|
| 94 |
+
- **Runtime:** Diffusers' built-in Nunchaku Lite quantizer for the DiT and the same `SVDQW4A4Linear` layers
|
| 95 |
+
for the text encoder, running the [nunchaku-lite](https://github.com/rootonchair/nunchaku-lite) CUDA kernels.
|
| 96 |
+
|
| 97 |
+
The scripts that produced this repository are in `tools/`: `quantize_dit.py`,
|
| 98 |
+
`quantize_text_encoder.py --rank 128 --edge-layers 4 --skip-attention`, `package.py`, `evaluate.py` and
|
| 99 |
+
`parrot_test.py`. They need a checkout of diffuse-compressor (its `examples/` package) at
|
| 100 |
+
`$DIFFUSE_COMPRESSOR`.
|
| 101 |
+
|
| 102 |
+
## Results
|
| 103 |
+
|
| 104 |
+
RTX PRO 6000 Blackwell (96 GB), Diffusers, 40 steps, guidance 1, no `torch.compile`, after a warmup.
|
| 105 |
+
Peak VRAM is PyTorch's peak allocation.
|
| 106 |
+
|
| 107 |
+
| | BF16 (official) | This repo |
|
| 108 |
+
|---|---|---|
|
| 109 |
+
| Text-to-image, 512x512 | 3.32 s | **1.91 s** (1.74x) |
|
| 110 |
+
| Edit, 512x512 | 3.73 s | **2.18 s** (1.71x) |
|
| 111 |
+
| Text-to-image, 1024x1024 | 14.06 s | **8.12 s** (1.73x) |
|
| 112 |
+
| Peak VRAM, 512x512 | 31.9-32.5 GiB | **18.6-19.3 GiB** |
|
| 113 |
+
| Throughput, 512x512 batch 1-8 | 0.28-0.30 img/s | **0.49-0.55 img/s** |
|
| 114 |
+
|
| 115 |
+
Batching several prompts gives no extra throughput in Diffusers on this GPU; it is saturated at batch 1.
|
| 116 |
+
The pipeline shares one reference-image list across a batch, so edits with different images run one at a time.
|
| 117 |
+
|
| 118 |
+
Closeness to BF16 (LPIPS with AlexNet, same seeds, 512x512; lower is closer, around 0.1 is hard to tell apart):
|
| 119 |
+
|
| 120 |
+
| Set | LPIPS vs BF16 |
|
| 121 |
+
|---|---|
|
| 122 |
+
| 8 text-to-image prompts (portraits, typography, counting, scenes) | 0.148 |
|
| 123 |
+
| 12 held-out NHR-Edit test edits + 1 parrot recolor | 0.046 |
|
| 124 |
+
|
| 125 |
+
Instruction following on a hard recolor ("Turn the parrots blue", 8 seeds, tiles with at least one blue parrot),
|
| 126 |
+
with this repo's DiT:
|
| 127 |
+
|
| 128 |
+
| Text encoder | Followed |
|
| 129 |
+
|---|---|
|
| 130 |
+
| BF16 | 7 / 8 |
|
| 131 |
+
| NVFP4, every linear, rank 32 | 1 / 8 |
|
| 132 |
+
| NVFP4, every linear, rank 128, 4 BF16 edge layers | 4-5 / 8 |
|
| 133 |
+
| **NVFP4, MLPs only, rank 128, 4 BF16 edge layers (this repo)** | **7 / 8** |
|
| 134 |
+
|
| 135 |
+
With the BF16 text encoder, this repo's DiT followed the edit in 8 of 8 seeds, the same as the BF16 DiT.
|
| 136 |
|
| 137 |
## Limitations
|
| 138 |
|
| 139 |
+
- Measured on one GPU model and one software stack; quantization shifts details of individual images.
|
| 140 |
+
- The evaluation is small (8 prompts, 13 edits, one 8-seed recolor test). It is not a benchmark.
|
| 141 |
+
- The kernel build covers `sm_120a`, PyTorch 2.13, CUDA 13 and CPython 3.12 only.
|
| 142 |
- Non-commercial research and evaluation use only.
|
| 143 |
+
|
| 144 |
+
## Previous version
|
| 145 |
+
|
| 146 |
+
Until 2026-09-25 this repository held a different combination for SGLang Diffusion: the ModelOpt NVFP4 DiT from
|
| 147 |
+
HangGlidersRule/Darkstar-Qwen-Image-2.1-Base-ModelOpt-W4A4-NVFP4 and the NVFP4 Qwen3-VL encoder from
|
| 148 |
+
BennyDaBall/Qwen-Image-2.1-NVFP4. It is still available in this repository's commit history. It was
|
| 149 |
+
replaced because its 4-bit text encoder lost hard edits and it did not load in Diffusers.
|
| 150 |
|
| 151 |
## Credits and license
|
| 152 |
|
| 153 |
+
- Qwen team: Qwen-Image-2.1.
|
| 154 |
+
- [SVDQuant](https://arxiv.org/abs/2411.05007) (MIT Han Lab), [nunchaku-lite](https://github.com/rootonchair/nunchaku-lite)
|
| 155 |
+
and [diffuse-compressor](https://github.com/rootonchair/diffuse-compressor) (rootonchair).
|
| 156 |
+
- Calibration edits: [VyoJ/NHR-Edit-Change_Only](https://huggingface.co/datasets/VyoJ/NHR-Edit-Change_Only).
|
| 157 |
|
| 158 |
+
Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory
|
| 159 |
+
Technology Co., Ltd. All Rights Reserved. See `LICENSE` and `NOTICE`, which lists every modified file.
|
model_index.json
CHANGED
|
@@ -10,15 +10,15 @@
|
|
| 10 |
"FlowMatchEulerDiscreteScheduler"
|
| 11 |
],
|
| 12 |
"text_encoder": [
|
| 13 |
-
"
|
| 14 |
-
"
|
| 15 |
],
|
| 16 |
"transformer": [
|
| 17 |
-
"
|
| 18 |
-
"
|
| 19 |
],
|
| 20 |
"vae": [
|
| 21 |
"diffusers",
|
| 22 |
"AutoencoderKLQwenImage21"
|
| 23 |
]
|
| 24 |
-
}
|
|
|
|
| 10 |
"FlowMatchEulerDiscreteScheduler"
|
| 11 |
],
|
| 12 |
"text_encoder": [
|
| 13 |
+
"modeling_nunchaku_qwen3vl",
|
| 14 |
+
"NunchakuQwen3VLForConditionalGeneration"
|
| 15 |
],
|
| 16 |
"transformer": [
|
| 17 |
+
"modeling_nunchaku_qwenimage21",
|
| 18 |
+
"NunchakuQwenImage21Transformer2DModel"
|
| 19 |
],
|
| 20 |
"vae": [
|
| 21 |
"diffusers",
|
| 22 |
"AutoencoderKLQwenImage21"
|
| 23 |
]
|
| 24 |
+
}
|
text_encoder/config.json
CHANGED
|
@@ -5,6 +5,101 @@
|
|
| 5 |
"dtype": "bfloat16",
|
| 6 |
"image_token_id": 151655,
|
| 7 |
"model_type": "qwen3_vl",
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
"text_config": {
|
| 9 |
"attention_bias": false,
|
| 10 |
"attention_dropout": 0.0,
|
|
|
|
| 5 |
"dtype": "bfloat16",
|
| 6 |
"image_token_id": 151655,
|
| 7 |
"model_type": "qwen3_vl",
|
| 8 |
+
"nunchaku_lite": {
|
| 9 |
+
"compute_dtype": "bfloat16",
|
| 10 |
+
"quant_method": "nunchaku_lite",
|
| 11 |
+
"svdq_w4a4": {
|
| 12 |
+
"group_size": 16,
|
| 13 |
+
"precision": "nvfp4",
|
| 14 |
+
"rank": 128,
|
| 15 |
+
"targets": [
|
| 16 |
+
"model.language_model.layers.4.mlp.gate_proj",
|
| 17 |
+
"model.language_model.layers.4.mlp.up_proj",
|
| 18 |
+
"model.language_model.layers.4.mlp.down_proj",
|
| 19 |
+
"model.language_model.layers.5.mlp.gate_proj",
|
| 20 |
+
"model.language_model.layers.5.mlp.up_proj",
|
| 21 |
+
"model.language_model.layers.5.mlp.down_proj",
|
| 22 |
+
"model.language_model.layers.6.mlp.gate_proj",
|
| 23 |
+
"model.language_model.layers.6.mlp.up_proj",
|
| 24 |
+
"model.language_model.layers.6.mlp.down_proj",
|
| 25 |
+
"model.language_model.layers.7.mlp.gate_proj",
|
| 26 |
+
"model.language_model.layers.7.mlp.up_proj",
|
| 27 |
+
"model.language_model.layers.7.mlp.down_proj",
|
| 28 |
+
"model.language_model.layers.8.mlp.gate_proj",
|
| 29 |
+
"model.language_model.layers.8.mlp.up_proj",
|
| 30 |
+
"model.language_model.layers.8.mlp.down_proj",
|
| 31 |
+
"model.language_model.layers.9.mlp.gate_proj",
|
| 32 |
+
"model.language_model.layers.9.mlp.up_proj",
|
| 33 |
+
"model.language_model.layers.9.mlp.down_proj",
|
| 34 |
+
"model.language_model.layers.10.mlp.gate_proj",
|
| 35 |
+
"model.language_model.layers.10.mlp.up_proj",
|
| 36 |
+
"model.language_model.layers.10.mlp.down_proj",
|
| 37 |
+
"model.language_model.layers.11.mlp.gate_proj",
|
| 38 |
+
"model.language_model.layers.11.mlp.up_proj",
|
| 39 |
+
"model.language_model.layers.11.mlp.down_proj",
|
| 40 |
+
"model.language_model.layers.12.mlp.gate_proj",
|
| 41 |
+
"model.language_model.layers.12.mlp.up_proj",
|
| 42 |
+
"model.language_model.layers.12.mlp.down_proj",
|
| 43 |
+
"model.language_model.layers.13.mlp.gate_proj",
|
| 44 |
+
"model.language_model.layers.13.mlp.up_proj",
|
| 45 |
+
"model.language_model.layers.13.mlp.down_proj",
|
| 46 |
+
"model.language_model.layers.14.mlp.gate_proj",
|
| 47 |
+
"model.language_model.layers.14.mlp.up_proj",
|
| 48 |
+
"model.language_model.layers.14.mlp.down_proj",
|
| 49 |
+
"model.language_model.layers.15.mlp.gate_proj",
|
| 50 |
+
"model.language_model.layers.15.mlp.up_proj",
|
| 51 |
+
"model.language_model.layers.15.mlp.down_proj",
|
| 52 |
+
"model.language_model.layers.16.mlp.gate_proj",
|
| 53 |
+
"model.language_model.layers.16.mlp.up_proj",
|
| 54 |
+
"model.language_model.layers.16.mlp.down_proj",
|
| 55 |
+
"model.language_model.layers.17.mlp.gate_proj",
|
| 56 |
+
"model.language_model.layers.17.mlp.up_proj",
|
| 57 |
+
"model.language_model.layers.17.mlp.down_proj",
|
| 58 |
+
"model.language_model.layers.18.mlp.gate_proj",
|
| 59 |
+
"model.language_model.layers.18.mlp.up_proj",
|
| 60 |
+
"model.language_model.layers.18.mlp.down_proj",
|
| 61 |
+
"model.language_model.layers.19.mlp.gate_proj",
|
| 62 |
+
"model.language_model.layers.19.mlp.up_proj",
|
| 63 |
+
"model.language_model.layers.19.mlp.down_proj",
|
| 64 |
+
"model.language_model.layers.20.mlp.gate_proj",
|
| 65 |
+
"model.language_model.layers.20.mlp.up_proj",
|
| 66 |
+
"model.language_model.layers.20.mlp.down_proj",
|
| 67 |
+
"model.language_model.layers.21.mlp.gate_proj",
|
| 68 |
+
"model.language_model.layers.21.mlp.up_proj",
|
| 69 |
+
"model.language_model.layers.21.mlp.down_proj",
|
| 70 |
+
"model.language_model.layers.22.mlp.gate_proj",
|
| 71 |
+
"model.language_model.layers.22.mlp.up_proj",
|
| 72 |
+
"model.language_model.layers.22.mlp.down_proj",
|
| 73 |
+
"model.language_model.layers.23.mlp.gate_proj",
|
| 74 |
+
"model.language_model.layers.23.mlp.up_proj",
|
| 75 |
+
"model.language_model.layers.23.mlp.down_proj",
|
| 76 |
+
"model.language_model.layers.24.mlp.gate_proj",
|
| 77 |
+
"model.language_model.layers.24.mlp.up_proj",
|
| 78 |
+
"model.language_model.layers.24.mlp.down_proj",
|
| 79 |
+
"model.language_model.layers.25.mlp.gate_proj",
|
| 80 |
+
"model.language_model.layers.25.mlp.up_proj",
|
| 81 |
+
"model.language_model.layers.25.mlp.down_proj",
|
| 82 |
+
"model.language_model.layers.26.mlp.gate_proj",
|
| 83 |
+
"model.language_model.layers.26.mlp.up_proj",
|
| 84 |
+
"model.language_model.layers.26.mlp.down_proj",
|
| 85 |
+
"model.language_model.layers.27.mlp.gate_proj",
|
| 86 |
+
"model.language_model.layers.27.mlp.up_proj",
|
| 87 |
+
"model.language_model.layers.27.mlp.down_proj",
|
| 88 |
+
"model.language_model.layers.28.mlp.gate_proj",
|
| 89 |
+
"model.language_model.layers.28.mlp.up_proj",
|
| 90 |
+
"model.language_model.layers.28.mlp.down_proj",
|
| 91 |
+
"model.language_model.layers.29.mlp.gate_proj",
|
| 92 |
+
"model.language_model.layers.29.mlp.up_proj",
|
| 93 |
+
"model.language_model.layers.29.mlp.down_proj",
|
| 94 |
+
"model.language_model.layers.30.mlp.gate_proj",
|
| 95 |
+
"model.language_model.layers.30.mlp.up_proj",
|
| 96 |
+
"model.language_model.layers.30.mlp.down_proj",
|
| 97 |
+
"model.language_model.layers.31.mlp.gate_proj",
|
| 98 |
+
"model.language_model.layers.31.mlp.up_proj",
|
| 99 |
+
"model.language_model.layers.31.mlp.down_proj"
|
| 100 |
+
]
|
| 101 |
+
}
|
| 102 |
+
},
|
| 103 |
"text_config": {
|
| 104 |
"attention_bias": false,
|
| 105 |
"attention_dropout": 0.0,
|
text_encoder/model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:83c4462b5ce6b98e2708ca53031e2f9a60a8e402c4c4a0ccd19cb2df8e817ae2
|
| 3 |
+
size 11812006400
|
text_encoder/modeling_nunchaku_qwen3vl.py
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Qwen3-VL text encoder with SVDQuant NVFP4 (Nunchaku Lite) linears.
|
| 2 |
+
|
| 3 |
+
Loaded by Diffusers as a custom pipeline component (`trust_remote_code=True`).
|
| 4 |
+
`config.json` holds the compact Nunchaku Lite config under `nunchaku_lite`; the
|
| 5 |
+
listed linears are swapped for Diffusers' own `SVDQW4A4Linear` before the
|
| 6 |
+
quantized state dict is loaded, so inference uses the same NVFP4 kernels as the
|
| 7 |
+
transformer. Everything else (vision tower, embeddings, the BF16 layers) loads
|
| 8 |
+
as in the stock model.
|
| 9 |
+
"""
|
| 10 |
+
|
| 11 |
+
from pathlib import Path
|
| 12 |
+
|
| 13 |
+
import torch
|
| 14 |
+
from accelerate import init_empty_weights
|
| 15 |
+
from huggingface_hub import snapshot_download
|
| 16 |
+
|
| 17 |
+
from .nunchaku_kernels import KERNELS # noqa: F401 (sets up the kernels before Diffusers imports them)
|
| 18 |
+
|
| 19 |
+
# isort: off -- must stay below `.nunchaku_kernels`, which prepares the kernels this import loads
|
| 20 |
+
from diffusers.quantizers.nunchaku.utils import replace_with_nunchaku_linear # noqa: E402
|
| 21 |
+
from safetensors.torch import load_file # noqa: E402
|
| 22 |
+
from transformers import Qwen3VLForConditionalGeneration # noqa: E402
|
| 23 |
+
# isort: on
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
class NunchakuQwen3VLForConditionalGeneration(Qwen3VLForConditionalGeneration):
|
| 27 |
+
@classmethod
|
| 28 |
+
def from_pretrained(cls, pretrained_model_name_or_path, *args, subfolder: str = "", **kwargs):
|
| 29 |
+
path = Path(pretrained_model_name_or_path) / subfolder
|
| 30 |
+
if not path.is_dir():
|
| 31 |
+
path = Path(snapshot_download(str(pretrained_model_name_or_path), allow_patterns=[f"{subfolder}/*"])) / subfolder
|
| 32 |
+
dtype = kwargs.get("dtype") or kwargs.get("torch_dtype") or torch.bfloat16
|
| 33 |
+
if dtype == "auto":
|
| 34 |
+
dtype = torch.bfloat16
|
| 35 |
+
|
| 36 |
+
config = cls.config_class.from_pretrained(path)
|
| 37 |
+
with init_empty_weights(): # parameters on meta, buffers (e.g. rotary tables) materialized
|
| 38 |
+
model = cls(config)
|
| 39 |
+
replace_with_nunchaku_linear(model, config.nunchaku_lite, dtype)
|
| 40 |
+
# strict=False only because non-persistent buffers are absent; both checks below stay strict
|
| 41 |
+
result = model.load_state_dict(load_file(path / "model.safetensors"), strict=False, assign=True)
|
| 42 |
+
if result.unexpected_keys:
|
| 43 |
+
raise ValueError(f"Checkpoint has {len(result.unexpected_keys)} unexpected tensors, e.g. {result.unexpected_keys[:3]}")
|
| 44 |
+
missing = [name for name, p in model.named_parameters() if p.device.type == "meta"]
|
| 45 |
+
if missing:
|
| 46 |
+
raise ValueError(f"Checkpoint is missing {len(missing)} parameters, e.g. {missing[:3]}")
|
| 47 |
+
return model.eval()
|
text_encoder/nunchaku_kernels.py
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Shared by text_encoder/ and transformer/: whichever component Diffusers loads first sets this up.
|
| 2 |
+
#
|
| 3 |
+
# Diffusers fetches the Nunchaku Lite kernels from `rootonchair/nunchaku-lite-kernels`, which is
|
| 4 |
+
# no longer downloadable. Loading this repo with trust_remote_code=True already runs this code,
|
| 5 |
+
# so point that kernel name at the rebuilt copy before Diffusers imports its Nunchaku utilities.
|
| 6 |
+
import os
|
| 7 |
+
|
| 8 |
+
from huggingface_hub import snapshot_download
|
| 9 |
+
|
| 10 |
+
KERNELS = "rootonchair/nunchaku-lite-kernels"
|
| 11 |
+
if KERNELS not in os.environ.get("LOCAL_KERNELS", ""):
|
| 12 |
+
local = f"{KERNELS}={snapshot_download('joseplcam/nunchaku-lite-kernels')}"
|
| 13 |
+
os.environ["LOCAL_KERNELS"] = ":".join(filter(None, [os.environ.get("LOCAL_KERNELS"), local]))
|
| 14 |
+
if "DIFFUSERS_TRUST_REMOTE_KERNELS" not in os.environ:
|
| 15 |
+
# Diffusers read this variable into a constant at import time, so set both.
|
| 16 |
+
import diffusers.utils.constants
|
| 17 |
+
|
| 18 |
+
os.environ["DIFFUSERS_TRUST_REMOTE_KERNELS"] = "true"
|
| 19 |
+
diffusers.utils.constants.DIFFUSERS_TRUST_REMOTE_KERNELS = True
|
tools/evaluate.py
ADDED
|
@@ -0,0 +1,103 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Compare a Qwen-Image-2.1 pipeline against the BF16 original on text-to-image and edits.
|
| 2 |
+
|
| 3 |
+
Writes, per pipeline: images, per-item LPIPS vs BF16 (same seeds), timings and peak VRAM.
|
| 4 |
+
Edits come from the NHR-Edit test split (held out from calibration) plus the parrot
|
| 5 |
+
recolor that 4-bit models failed before. Instruction following is judged by eye from
|
| 6 |
+
the saved grids; LPIPS only measures closeness to BF16.
|
| 7 |
+
|
| 8 |
+
Usage:
|
| 9 |
+
python tools/qwen21_nvfp4/evaluate.py <pipeline path or id> <tag> [--trust-remote-code] [--baseline <bf16 out dir>]
|
| 10 |
+
"""
|
| 11 |
+
|
| 12 |
+
import argparse
|
| 13 |
+
import json
|
| 14 |
+
import time
|
| 15 |
+
import urllib.request
|
| 16 |
+
from io import BytesIO
|
| 17 |
+
from pathlib import Path
|
| 18 |
+
|
| 19 |
+
import datasets
|
| 20 |
+
import torch
|
| 21 |
+
from diffusers import DiffusionPipeline
|
| 22 |
+
from PIL import Image
|
| 23 |
+
|
| 24 |
+
PROMPTS = [
|
| 25 |
+
"A capybara reading a book by candlelight",
|
| 26 |
+
"A red vintage bicycle leaning against a blue door",
|
| 27 |
+
"A bowl of ramen with a soft boiled egg, top view",
|
| 28 |
+
"A neon sign that says OPEN in a rainy alley",
|
| 29 |
+
"A glass teapot with blooming flower tea on a table",
|
| 30 |
+
"A portrait of an elderly fisherman with a knitted cap, soft window light",
|
| 31 |
+
'A chalkboard menu that reads "SOUP OF THE DAY: TOMATO"',
|
| 32 |
+
"Three red apples and two green pears on a wooden table",
|
| 33 |
+
]
|
| 34 |
+
PARROTS = "https://r0k.us/graphics/kodak/kodak/kodim23.png"
|
| 35 |
+
|
| 36 |
+
|
| 37 |
+
def edit_items(n: int) -> list[tuple[str, Image.Image, str]]:
|
| 38 |
+
data = datasets.load_dataset("VyoJ/NHR-Edit-Change_Only", split="test")
|
| 39 |
+
items = [(f"nhr{i}", data[i]["source"].convert("RGB"), data[i]["edit_instruction"]) for i in range(n)]
|
| 40 |
+
parrots = Image.open(BytesIO(urllib.request.urlopen(PARROTS).read())).convert("RGB")
|
| 41 |
+
return items + [("parrots", parrots, "Turn the parrots blue")]
|
| 42 |
+
|
| 43 |
+
|
| 44 |
+
def main():
|
| 45 |
+
parser = argparse.ArgumentParser()
|
| 46 |
+
parser.add_argument("pipeline")
|
| 47 |
+
parser.add_argument("tag")
|
| 48 |
+
parser.add_argument("--trust-remote-code", action="store_true")
|
| 49 |
+
parser.add_argument("--edits", type=int, default=12)
|
| 50 |
+
parser.add_argument("--size", type=int, default=512)
|
| 51 |
+
parser.add_argument("--out", default="outputs/qwen21_eval")
|
| 52 |
+
parser.add_argument("--baseline", help="output dir of the BF16 run, for LPIPS")
|
| 53 |
+
args = parser.parse_args()
|
| 54 |
+
|
| 55 |
+
out = Path(args.out) / args.tag
|
| 56 |
+
out.mkdir(parents=True, exist_ok=True)
|
| 57 |
+
pipe = DiffusionPipeline.from_pretrained(args.pipeline, dtype=torch.bfloat16, trust_remote_code=args.trust_remote_code).to("cuda")
|
| 58 |
+
pipe.set_progress_bar_config(disable=True)
|
| 59 |
+
gen = lambda seed: torch.Generator("cuda").manual_seed(seed) # noqa: E731
|
| 60 |
+
|
| 61 |
+
def run(name, seed, **kwargs):
|
| 62 |
+
torch.cuda.synchronize()
|
| 63 |
+
start = time.perf_counter()
|
| 64 |
+
image = pipe(num_inference_steps=40, generator=gen(seed), **kwargs).images[0]
|
| 65 |
+
torch.cuda.synchronize()
|
| 66 |
+
image.save(out / f"{name}.png")
|
| 67 |
+
return time.perf_counter() - start
|
| 68 |
+
|
| 69 |
+
run("warmup", 0, prompt="warmup", height=args.size, width=args.size)
|
| 70 |
+
run("warmup", 1, prompt="warmup", image=Image.open(out / "warmup.png"), output_resolution=args.size)
|
| 71 |
+
torch.cuda.reset_peak_memory_stats()
|
| 72 |
+
times = {"t2i": [], "edit": []}
|
| 73 |
+
for i, prompt in enumerate(PROMPTS):
|
| 74 |
+
times["t2i"].append(run(f"t2i{i}", 1000 + i, prompt=prompt, height=args.size, width=args.size))
|
| 75 |
+
for name, image, prompt in edit_items(args.edits):
|
| 76 |
+
times["edit"].append(run(name, 2000, prompt=prompt, image=image, output_resolution=args.size))
|
| 77 |
+
report = {
|
| 78 |
+
"tag": args.tag,
|
| 79 |
+
"size": args.size,
|
| 80 |
+
"t2i_s": round(sum(times["t2i"]) / len(times["t2i"]), 3),
|
| 81 |
+
"edit_s": round(sum(times["edit"]) / len(times["edit"]), 3),
|
| 82 |
+
"peak_vram_gib": round(torch.cuda.max_memory_allocated() / 2**30, 2),
|
| 83 |
+
}
|
| 84 |
+
if args.baseline:
|
| 85 |
+
import lpips
|
| 86 |
+
import numpy as np
|
| 87 |
+
|
| 88 |
+
metric = lpips.LPIPS(net="alex", verbose=False)
|
| 89 |
+
load = lambda p: torch.from_numpy(np.array(Image.open(p).convert("RGB"))).permute(2, 0, 1)[None].float() / 127.5 - 1 # noqa: E731
|
| 90 |
+
scores = {
|
| 91 |
+
p.stem: round(metric(load(Path(args.baseline) / p.name), load(p)).item(), 4)
|
| 92 |
+
for p in sorted(out.glob("*.png"))
|
| 93 |
+
if p.stem != "warmup" and (Path(args.baseline) / p.name).exists()
|
| 94 |
+
}
|
| 95 |
+
report["lpips"] = scores
|
| 96 |
+
report["lpips_t2i_mean"] = round(sum(v for k, v in scores.items() if k.startswith("t2i")) / len(PROMPTS), 4)
|
| 97 |
+
report["lpips_edit_mean"] = round(sum(v for k, v in scores.items() if not k.startswith("t2i")) / (args.edits + 1), 4)
|
| 98 |
+
(out / "report.json").write_text(json.dumps(report, indent=1))
|
| 99 |
+
print(json.dumps(report))
|
| 100 |
+
|
| 101 |
+
|
| 102 |
+
if __name__ == "__main__":
|
| 103 |
+
main()
|
tools/modeling_nunchaku_qwen3vl.py
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Qwen3-VL text encoder with SVDQuant NVFP4 (Nunchaku Lite) linears.
|
| 2 |
+
|
| 3 |
+
Loaded by Diffusers as a custom pipeline component (`trust_remote_code=True`).
|
| 4 |
+
`config.json` holds the compact Nunchaku Lite config under `nunchaku_lite`; the
|
| 5 |
+
listed linears are swapped for Diffusers' own `SVDQW4A4Linear` before the
|
| 6 |
+
quantized state dict is loaded, so inference uses the same NVFP4 kernels as the
|
| 7 |
+
transformer. Everything else (vision tower, embeddings, the BF16 layers) loads
|
| 8 |
+
as in the stock model.
|
| 9 |
+
"""
|
| 10 |
+
|
| 11 |
+
from pathlib import Path
|
| 12 |
+
|
| 13 |
+
import torch
|
| 14 |
+
from accelerate import init_empty_weights
|
| 15 |
+
from huggingface_hub import snapshot_download
|
| 16 |
+
|
| 17 |
+
from .nunchaku_kernels import KERNELS # noqa: F401 (sets up the kernels before Diffusers imports them)
|
| 18 |
+
|
| 19 |
+
# isort: off -- must stay below `.nunchaku_kernels`, which prepares the kernels this import loads
|
| 20 |
+
from diffusers.quantizers.nunchaku.utils import replace_with_nunchaku_linear # noqa: E402
|
| 21 |
+
from safetensors.torch import load_file # noqa: E402
|
| 22 |
+
from transformers import Qwen3VLForConditionalGeneration # noqa: E402
|
| 23 |
+
# isort: on
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
class NunchakuQwen3VLForConditionalGeneration(Qwen3VLForConditionalGeneration):
|
| 27 |
+
@classmethod
|
| 28 |
+
def from_pretrained(cls, pretrained_model_name_or_path, *args, subfolder: str = "", **kwargs):
|
| 29 |
+
path = Path(pretrained_model_name_or_path) / subfolder
|
| 30 |
+
if not path.is_dir():
|
| 31 |
+
path = Path(snapshot_download(str(pretrained_model_name_or_path), allow_patterns=[f"{subfolder}/*"])) / subfolder
|
| 32 |
+
dtype = kwargs.get("dtype") or kwargs.get("torch_dtype") or torch.bfloat16
|
| 33 |
+
if dtype == "auto":
|
| 34 |
+
dtype = torch.bfloat16
|
| 35 |
+
|
| 36 |
+
config = cls.config_class.from_pretrained(path)
|
| 37 |
+
with init_empty_weights(): # parameters on meta, buffers (e.g. rotary tables) materialized
|
| 38 |
+
model = cls(config)
|
| 39 |
+
replace_with_nunchaku_linear(model, config.nunchaku_lite, dtype)
|
| 40 |
+
# strict=False only because non-persistent buffers are absent; both checks below stay strict
|
| 41 |
+
result = model.load_state_dict(load_file(path / "model.safetensors"), strict=False, assign=True)
|
| 42 |
+
if result.unexpected_keys:
|
| 43 |
+
raise ValueError(f"Checkpoint has {len(result.unexpected_keys)} unexpected tensors, e.g. {result.unexpected_keys[:3]}")
|
| 44 |
+
missing = [name for name, p in model.named_parameters() if p.device.type == "meta"]
|
| 45 |
+
if missing:
|
| 46 |
+
raise ValueError(f"Checkpoint is missing {len(missing)} parameters, e.g. {missing[:3]}")
|
| 47 |
+
return model.eval()
|
tools/modeling_nunchaku_qwenimage21.py
ADDED
|
@@ -0,0 +1,14 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Qwen-Image-2.1 transformer, unchanged, loaded as a custom component.
|
| 2 |
+
|
| 3 |
+
Its only job is importing `nunchaku_kernels`, so the NVFP4 kernels are set up even when
|
| 4 |
+
Diffusers loads this component before the text encoder. The Nunchaku Lite quantization
|
| 5 |
+
itself is handled by Diffusers' built-in quantizer from `config.json`.
|
| 6 |
+
"""
|
| 7 |
+
|
| 8 |
+
from diffusers import QwenImage21Transformer2DModel
|
| 9 |
+
|
| 10 |
+
from .nunchaku_kernels import KERNELS # noqa: F401
|
| 11 |
+
|
| 12 |
+
|
| 13 |
+
class NunchakuQwenImage21Transformer2DModel(QwenImage21Transformer2DModel):
|
| 14 |
+
pass
|
tools/nunchaku_kernels.py
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Shared by text_encoder/ and transformer/: whichever component Diffusers loads first sets this up.
|
| 2 |
+
#
|
| 3 |
+
# Diffusers fetches the Nunchaku Lite kernels from `rootonchair/nunchaku-lite-kernels`, which is
|
| 4 |
+
# no longer downloadable. Loading this repo with trust_remote_code=True already runs this code,
|
| 5 |
+
# so point that kernel name at the rebuilt copy before Diffusers imports its Nunchaku utilities.
|
| 6 |
+
import os
|
| 7 |
+
|
| 8 |
+
from huggingface_hub import snapshot_download
|
| 9 |
+
|
| 10 |
+
KERNELS = "rootonchair/nunchaku-lite-kernels"
|
| 11 |
+
if KERNELS not in os.environ.get("LOCAL_KERNELS", ""):
|
| 12 |
+
local = f"{KERNELS}={snapshot_download('joseplcam/nunchaku-lite-kernels')}"
|
| 13 |
+
os.environ["LOCAL_KERNELS"] = ":".join(filter(None, [os.environ.get("LOCAL_KERNELS"), local]))
|
| 14 |
+
if "DIFFUSERS_TRUST_REMOTE_KERNELS" not in os.environ:
|
| 15 |
+
# Diffusers read this variable into a constant at import time, so set both.
|
| 16 |
+
import diffusers.utils.constants
|
| 17 |
+
|
| 18 |
+
os.environ["DIFFUSERS_TRUST_REMOTE_KERNELS"] = "true"
|
| 19 |
+
diffusers.utils.constants.DIFFUSERS_TRUST_REMOTE_KERNELS = True
|
tools/package.py
ADDED
|
@@ -0,0 +1,81 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Package the quantized DiT and text encoder into one Diffusers pipeline folder.
|
| 2 |
+
|
| 3 |
+
- transformer/: Nunchaku Lite NVFP4 checkpoint (loaded by Diffusers' built-in quantizer)
|
| 4 |
+
- text_encoder/: Nunchaku Lite NVFP4 checkpoint + custom component class (trust_remote_code)
|
| 5 |
+
- vae/: official weights stored in BF16
|
| 6 |
+
- processor/, scheduler/, model_index.json, LICENSE: from Qwen/Qwen-Image-2.1
|
| 7 |
+
|
| 8 |
+
Usage:
|
| 9 |
+
python tools/qwen21_nvfp4/package.py --dit <dit.safetensors> --text-encoder <te.safetensors> --out <dir>
|
| 10 |
+
"""
|
| 11 |
+
|
| 12 |
+
import argparse
|
| 13 |
+
import json
|
| 14 |
+
import os
|
| 15 |
+
import shutil
|
| 16 |
+
import sys
|
| 17 |
+
from pathlib import Path
|
| 18 |
+
|
| 19 |
+
import torch
|
| 20 |
+
from diffusers import AutoencoderKLQwenImage21
|
| 21 |
+
|
| 22 |
+
# examples/ is not part of the installed package, so a checkout of the repo is needed
|
| 23 |
+
sys.path.insert(0, os.environ.get("DIFFUSE_COMPRESSOR", str(Path.home() / "src" / "diffuse-compressor")))
|
| 24 |
+
from examples.convert_nunchaku_lite_diffusers import ( # noqa: E402
|
| 25 |
+
build_diffusers_quantization_config,
|
| 26 |
+
package_diffusers_pipeline,
|
| 27 |
+
)
|
| 28 |
+
|
| 29 |
+
HERE = Path(__file__).parent
|
| 30 |
+
# component -> (module file, class); both import nunchaku_kernels, which ships next to each
|
| 31 |
+
CUSTOM_COMPONENTS = {
|
| 32 |
+
"text_encoder": ("modeling_nunchaku_qwen3vl", "NunchakuQwen3VLForConditionalGeneration"),
|
| 33 |
+
"transformer": ("modeling_nunchaku_qwenimage21", "NunchakuQwenImage21Transformer2DModel"),
|
| 34 |
+
}
|
| 35 |
+
|
| 36 |
+
|
| 37 |
+
def main():
|
| 38 |
+
parser = argparse.ArgumentParser()
|
| 39 |
+
parser.add_argument("--dit", required=True)
|
| 40 |
+
parser.add_argument("--text-encoder", required=True)
|
| 41 |
+
parser.add_argument("--out", required=True)
|
| 42 |
+
parser.add_argument("--model-id", default="Qwen/Qwen-Image-2.1")
|
| 43 |
+
args = parser.parse_args()
|
| 44 |
+
|
| 45 |
+
out = package_diffusers_pipeline(args.dit, args.model_id, args.out, compute_dtype="bfloat16")
|
| 46 |
+
|
| 47 |
+
te = out / "text_encoder"
|
| 48 |
+
for dense in [*te.glob("*.safetensors"), *te.glob("*.safetensors.index.json")]:
|
| 49 |
+
dense.unlink()
|
| 50 |
+
shutil.copy2(args.text_encoder, te / "model.safetensors")
|
| 51 |
+
config = json.loads((te / "config.json").read_text())
|
| 52 |
+
config["nunchaku_lite"] = build_diffusers_quantization_config(args.text_encoder, compute_dtype="bfloat16")
|
| 53 |
+
(te / "config.json").write_text(json.dumps(config, indent=2, sort_keys=True) + "\n")
|
| 54 |
+
|
| 55 |
+
index = json.loads((out / "model_index.json").read_text())
|
| 56 |
+
for component, (module, cls) in CUSTOM_COMPONENTS.items():
|
| 57 |
+
shutil.copy2(HERE / f"{module}.py", out / component / f"{module}.py")
|
| 58 |
+
shutil.copy2(HERE / "nunchaku_kernels.py", out / component / "nunchaku_kernels.py")
|
| 59 |
+
index[component] = [module, cls]
|
| 60 |
+
(out / "model_index.json").write_text(json.dumps(index, indent=2) + "\n")
|
| 61 |
+
|
| 62 |
+
vae = AutoencoderKLQwenImage21.from_pretrained(out / "vae", dtype=torch.bfloat16)
|
| 63 |
+
shutil.rmtree(out / "vae")
|
| 64 |
+
vae.save_pretrained(out / "vae")
|
| 65 |
+
vae_config = json.loads((out / "vae" / "config.json").read_text())
|
| 66 |
+
vae_config.pop("_name_or_path", None) # a local path, meaningless on the Hub
|
| 67 |
+
(out / "vae" / "config.json").write_text(json.dumps(vae_config, indent=2, sort_keys=True) + "\n")
|
| 68 |
+
|
| 69 |
+
for leftover in ("assets", ".gitattributes"):
|
| 70 |
+
path = out / leftover
|
| 71 |
+
shutil.rmtree(path) if path.is_dir() else path.unlink(missing_ok=True)
|
| 72 |
+
for doc in ("README.md", "NOTICE"): # model card and license notice
|
| 73 |
+
shutil.copy2(HERE / doc, out / doc)
|
| 74 |
+
(out / "tools").mkdir()
|
| 75 |
+
for script in sorted(HERE.glob("*.py")): # the recipe, for reproducibility
|
| 76 |
+
shutil.copy2(script, out / "tools" / script.name)
|
| 77 |
+
print(out)
|
| 78 |
+
|
| 79 |
+
|
| 80 |
+
if __name__ == "__main__":
|
| 81 |
+
main()
|
tools/parrot_test.py
ADDED
|
@@ -0,0 +1,69 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Hard-edit check: "Turn the parrots blue" over 8 seeds at 512x512, per text encoder variant.
|
| 2 |
+
|
| 3 |
+
4-bit text encoders tend to lose this recolor (the model keeps the reference colors), so it
|
| 4 |
+
separates encoder variants that LPIPS alone does not. Each row of the output grid is one
|
| 5 |
+
variant; count the tiles with at least one blue parrot.
|
| 6 |
+
|
| 7 |
+
Usage:
|
| 8 |
+
python tools/qwen21_nvfp4/parrot_test.py <packaged pipeline> <out.jpg> bf16 <encoder checkpoint> ...
|
| 9 |
+
"""
|
| 10 |
+
|
| 11 |
+
import io
|
| 12 |
+
import json
|
| 13 |
+
import shutil
|
| 14 |
+
import sys
|
| 15 |
+
import tempfile
|
| 16 |
+
import urllib.request
|
| 17 |
+
from pathlib import Path
|
| 18 |
+
|
| 19 |
+
import torch
|
| 20 |
+
from diffusers import DiffusionPipeline
|
| 21 |
+
from PIL import Image
|
| 22 |
+
from transformers import Qwen3VLForConditionalGeneration
|
| 23 |
+
|
| 24 |
+
sys.path.insert(0, str(Path(__file__).parent))
|
| 25 |
+
from package import build_diffusers_quantization_config # noqa: E402
|
| 26 |
+
|
| 27 |
+
PARROTS = "https://r0k.us/graphics/kodak/kodak/kodim23.png"
|
| 28 |
+
|
| 29 |
+
|
| 30 |
+
def load_encoder(encoder_cls, pipeline_dir: Path, variant: str):
|
| 31 |
+
if variant == "bf16":
|
| 32 |
+
return Qwen3VLForConditionalGeneration.from_pretrained("Qwen/Qwen-Image-2.1", subfolder="text_encoder", dtype=torch.bfloat16)
|
| 33 |
+
with tempfile.TemporaryDirectory() as tmp: # the packaged text_encoder folder with this checkpoint swapped in
|
| 34 |
+
folder = Path(tmp)
|
| 35 |
+
for name in ("generation_config.json",):
|
| 36 |
+
shutil.copy2(pipeline_dir / "text_encoder" / name, folder / name)
|
| 37 |
+
config = json.loads((pipeline_dir / "text_encoder" / "config.json").read_text())
|
| 38 |
+
config["nunchaku_lite"] = build_diffusers_quantization_config(variant, compute_dtype="bfloat16")
|
| 39 |
+
(folder / "config.json").write_text(json.dumps(config))
|
| 40 |
+
(folder / "model.safetensors").symlink_to(Path(variant).resolve())
|
| 41 |
+
return encoder_cls.from_pretrained(folder, dtype=torch.bfloat16)
|
| 42 |
+
|
| 43 |
+
|
| 44 |
+
def main():
|
| 45 |
+
pipeline_dir, out, variants = Path(sys.argv[1]), sys.argv[2], sys.argv[3:]
|
| 46 |
+
pipe = DiffusionPipeline.from_pretrained(pipeline_dir, dtype=torch.bfloat16, trust_remote_code=True).to("cuda")
|
| 47 |
+
pipe.set_progress_bar_config(disable=True)
|
| 48 |
+
parrots = Image.open(io.BytesIO(urllib.request.urlopen(PARROTS).read())).convert("RGB")
|
| 49 |
+
encoder_cls = type(pipe.text_encoder)
|
| 50 |
+
rows = []
|
| 51 |
+
for variant in variants:
|
| 52 |
+
pipe.text_encoder = None
|
| 53 |
+
torch.cuda.empty_cache()
|
| 54 |
+
pipe.text_encoder = load_encoder(encoder_cls, pipeline_dir, variant).to("cuda")
|
| 55 |
+
rows.append([
|
| 56 |
+
pipe(prompt="Turn the parrots blue", image=parrots, output_resolution=512, num_inference_steps=40,
|
| 57 |
+
generator=torch.Generator("cuda").manual_seed(seed)).images[0].convert("RGB").resize((128, 128))
|
| 58 |
+
for seed in range(1, 9)
|
| 59 |
+
])
|
| 60 |
+
print(Path(variant).name, "done", flush=True)
|
| 61 |
+
grid = Image.new("RGB", (128 * 8, 128 * len(rows)))
|
| 62 |
+
for y, row in enumerate(rows):
|
| 63 |
+
for x, tile in enumerate(row):
|
| 64 |
+
grid.paste(tile, (128 * x, 128 * y))
|
| 65 |
+
grid.save(out)
|
| 66 |
+
|
| 67 |
+
|
| 68 |
+
if __name__ == "__main__":
|
| 69 |
+
main()
|
tools/quantize_dit.py
ADDED
|
@@ -0,0 +1,144 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""SVDQuant + GPTQ NVFP4 quantization of the Qwen-Image-2.1 DiT with diffuse-compressor.
|
| 2 |
+
|
| 3 |
+
Calibration mixes text-to-image prompts and image edits (NHR-Edit train split),
|
| 4 |
+
mostly at 512x512 with some 1024x1024, so both workloads shape the scales.
|
| 5 |
+
The first two and last two blocks and the global modulation stay in BF16.
|
| 6 |
+
|
| 7 |
+
Usage (from the repo root, envs/base_final interpreter):
|
| 8 |
+
python tools/qwen21_nvfp4/quantize_dit.py --num-samples 128 --rank 32
|
| 9 |
+
"""
|
| 10 |
+
|
| 11 |
+
import argparse
|
| 12 |
+
import os
|
| 13 |
+
import random
|
| 14 |
+
import sys
|
| 15 |
+
from dataclasses import replace
|
| 16 |
+
from pathlib import Path
|
| 17 |
+
|
| 18 |
+
import torch
|
| 19 |
+
from diffusers import QwenImage21Pipeline
|
| 20 |
+
|
| 21 |
+
# examples/ is not part of the installed package, so a checkout of the repo is needed
|
| 22 |
+
sys.path.insert(0, os.environ.get("DIFFUSE_COMPRESSOR", str(Path.home() / "src" / "diffuse-compressor")))
|
| 23 |
+
|
| 24 |
+
from diffuse_compressor import ( # noqa: E402
|
| 25 |
+
CalibrationSpec,
|
| 26 |
+
ExportSpec,
|
| 27 |
+
LoggingConfig,
|
| 28 |
+
QuantizationCacheSpec,
|
| 29 |
+
quantize_and_export,
|
| 30 |
+
)
|
| 31 |
+
from diffuse_compressor.config import GptqSpec # noqa: E402
|
| 32 |
+
from examples.image_to_image.quantize_hf import scan_linear_targets # noqa: E402
|
| 33 |
+
from examples.text_to_image.utils import ( # noqa: E402
|
| 34 |
+
make_generator,
|
| 35 |
+
save_diffusers_images,
|
| 36 |
+
standard_prompt_records,
|
| 37 |
+
svdquant_spec,
|
| 38 |
+
)
|
| 39 |
+
|
| 40 |
+
SKIP = [
|
| 41 |
+
"transformer_blocks.0.*",
|
| 42 |
+
"transformer_blocks.1.*",
|
| 43 |
+
"transformer_blocks.30.*",
|
| 44 |
+
"transformer_blocks.31.*",
|
| 45 |
+
"modulation*",
|
| 46 |
+
]
|
| 47 |
+
|
| 48 |
+
|
| 49 |
+
def records(num_samples: int, edit_dataset: str) -> list[dict]:
|
| 50 |
+
import datasets
|
| 51 |
+
|
| 52 |
+
half = num_samples // 2
|
| 53 |
+
rng = random.Random(0)
|
| 54 |
+
# 3 of every 4 samples at 512x512 (the working resolution), the rest at 1024x1024
|
| 55 |
+
size = lambda i: 1024 if i % 4 == 3 else 512 # noqa: E731
|
| 56 |
+
|
| 57 |
+
t2i = [
|
| 58 |
+
{"filename": r["filename"], "prompt": r["prompt"], "seed": r["seed"], "size": size(i)}
|
| 59 |
+
for i, r in enumerate(standard_prompt_records(num_samples - half))
|
| 60 |
+
]
|
| 61 |
+
data = datasets.load_dataset(edit_dataset, split="train")
|
| 62 |
+
edits = [
|
| 63 |
+
{
|
| 64 |
+
"filename": f"edit-{j}",
|
| 65 |
+
"prompt": data[j]["edit_instruction"],
|
| 66 |
+
"image": data[j]["source"].convert("RGB"),
|
| 67 |
+
"seed": 100_000 + j,
|
| 68 |
+
"size": size(i),
|
| 69 |
+
}
|
| 70 |
+
for i, j in enumerate(sorted(rng.sample(range(len(data)), half)))
|
| 71 |
+
]
|
| 72 |
+
return t2i + edits
|
| 73 |
+
|
| 74 |
+
|
| 75 |
+
def forward_fn(pipe, steps: int, device: str):
|
| 76 |
+
def forward(sample: dict):
|
| 77 |
+
kwargs = {
|
| 78 |
+
"prompt": sample["prompt"],
|
| 79 |
+
"num_inference_steps": steps,
|
| 80 |
+
"generator": make_generator(sample["seed"], device=device),
|
| 81 |
+
# the prefix KV cache is a stateful object that calibration replay cannot move or
|
| 82 |
+
# re-run; without it the prefix tokens pass through the same linears every step
|
| 83 |
+
"use_kv_cache": False,
|
| 84 |
+
}
|
| 85 |
+
if sample.get("image") is not None:
|
| 86 |
+
kwargs.update(image=sample["image"], output_resolution=sample["size"])
|
| 87 |
+
else:
|
| 88 |
+
kwargs.update(height=sample["size"], width=sample["size"])
|
| 89 |
+
return pipe(**kwargs)
|
| 90 |
+
|
| 91 |
+
return forward
|
| 92 |
+
|
| 93 |
+
|
| 94 |
+
def main():
|
| 95 |
+
parser = argparse.ArgumentParser()
|
| 96 |
+
parser.add_argument("--model-id", default="Qwen/Qwen-Image-2.1")
|
| 97 |
+
parser.add_argument("--edit-dataset", default="VyoJ/NHR-Edit-Change_Only")
|
| 98 |
+
parser.add_argument("--num-samples", type=int, default=128)
|
| 99 |
+
parser.add_argument("--steps", type=int, default=20)
|
| 100 |
+
parser.add_argument("--rank", type=int, default=32)
|
| 101 |
+
parser.add_argument("--no-gptq", action="store_true")
|
| 102 |
+
parser.add_argument("--out-dir", default="outputs/qwen21_nvfp4")
|
| 103 |
+
parser.add_argument("--cache-mode", choices=("reuse", "refresh", "disabled"), default="reuse")
|
| 104 |
+
parser.add_argument("--inspect", action="store_true")
|
| 105 |
+
args = parser.parse_args()
|
| 106 |
+
|
| 107 |
+
device = "cuda"
|
| 108 |
+
pipe = QwenImage21Pipeline.from_pretrained(args.model_id, torch_dtype=torch.bfloat16).to(device)
|
| 109 |
+
pipe.set_progress_bar_config(disable=True)
|
| 110 |
+
scan = scan_linear_targets(pipe.transformer, precision="nvfp4", rank=args.rank, skip=SKIP)
|
| 111 |
+
print(scan.format_text().splitlines()[0])
|
| 112 |
+
if args.inspect:
|
| 113 |
+
print(scan.format_text())
|
| 114 |
+
return
|
| 115 |
+
|
| 116 |
+
tag = f"svdq-nvfp4_r{args.rank}{'' if args.no_gptq else '-gptq'}-qwen-image-2.1"
|
| 117 |
+
out = Path(args.out_dir)
|
| 118 |
+
cache = out / "calibration" / tag
|
| 119 |
+
spec = replace(svdquant_spec("nvfp4", compute_device=device), rank=args.rank, gptq=GptqSpec(enabled=not args.no_gptq))
|
| 120 |
+
quantize_and_export(
|
| 121 |
+
pipe.transformer,
|
| 122 |
+
spec,
|
| 123 |
+
scan.target_config,
|
| 124 |
+
CalibrationSpec(
|
| 125 |
+
samples=records(args.num_samples, args.edit_dataset),
|
| 126 |
+
num_samples=args.num_samples,
|
| 127 |
+
cache_num_samples=args.num_samples,
|
| 128 |
+
batch_size=1,
|
| 129 |
+
cache_dir=cache / "inputs",
|
| 130 |
+
cache_mode=args.cache_mode,
|
| 131 |
+
forward_fn=forward_fn(pipe, args.steps, device),
|
| 132 |
+
max_rows_per_target=4096,
|
| 133 |
+
artifact_cache=QuantizationCacheSpec(cache / "artifacts", args.cache_mode),
|
| 134 |
+
output_dir=cache / "inputs" / "samples",
|
| 135 |
+
output_save_fn=save_diffusers_images,
|
| 136 |
+
scope_capture_mode="all_targets",
|
| 137 |
+
),
|
| 138 |
+
ExportSpec(output=out / "checkpoints" / f"{tag}.safetensors"),
|
| 139 |
+
LoggingConfig(name=tag),
|
| 140 |
+
)
|
| 141 |
+
|
| 142 |
+
|
| 143 |
+
if __name__ == "__main__":
|
| 144 |
+
main()
|
tools/quantize_text_encoder.py
ADDED
|
@@ -0,0 +1,121 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""SVDQuant + GPTQ NVFP4 quantization of the Qwen-Image-2.1 text encoder (Qwen3-VL 8B).
|
| 2 |
+
|
| 3 |
+
Only the language-model linears are quantized. The vision tower, embeddings,
|
| 4 |
+
lm_head (unused by the pipeline) and the first/last `--edge-layers` decoder
|
| 5 |
+
layers stay BF16: the pipeline feeds the last decoder layer's output straight
|
| 6 |
+
into the DiT. Calibration uses the same text-to-image and edit records as the DiT, run
|
| 7 |
+
through the full pipeline so condition images are resized exactly as in use.
|
| 8 |
+
|
| 9 |
+
Usage (from the repo root, envs/base_final interpreter):
|
| 10 |
+
python tools/qwen21_nvfp4/quantize_text_encoder.py --num-samples 128 --rank 32
|
| 11 |
+
"""
|
| 12 |
+
|
| 13 |
+
import argparse
|
| 14 |
+
import sys
|
| 15 |
+
from dataclasses import replace
|
| 16 |
+
from pathlib import Path
|
| 17 |
+
|
| 18 |
+
import torch
|
| 19 |
+
from diffusers import QwenImage21Pipeline
|
| 20 |
+
|
| 21 |
+
sys.path.insert(0, str(Path(__file__).parent))
|
| 22 |
+
from quantize_dit import records # noqa: E402 (also puts diffuse-compressor on sys.path)
|
| 23 |
+
|
| 24 |
+
from diffuse_compressor import ( # noqa: E402
|
| 25 |
+
CalibrationSpec,
|
| 26 |
+
ExportSpec,
|
| 27 |
+
LoggingConfig,
|
| 28 |
+
QuantizationCacheSpec,
|
| 29 |
+
quantize_and_export,
|
| 30 |
+
)
|
| 31 |
+
from diffuse_compressor.calibration import replay # noqa: E402
|
| 32 |
+
from diffuse_compressor.config import GptqSpec # noqa: E402
|
| 33 |
+
from examples.image_to_image.quantize_hf import scan_linear_targets # noqa: E402
|
| 34 |
+
from examples.text_to_image.utils import make_generator, svdquant_spec # noqa: E402
|
| 35 |
+
|
| 36 |
+
# Full-model replay CPU-offloads the model through Accelerate, which puts the 16 GB encoder in
|
| 37 |
+
# host RAM and trips the 90% RAM guard on a 32 GB machine. The GPU has room for it, and
|
| 38 |
+
# returning False makes the library fall back to `model.to(device)`.
|
| 39 |
+
replay._accelerate_cpu_offload_for_full_replay = lambda *args, **kwargs: False
|
| 40 |
+
|
| 41 |
+
NUM_LAYERS = 36
|
| 42 |
+
|
| 43 |
+
|
| 44 |
+
def skip_patterns(edge_layers: int, skip_attention: bool) -> list[str]:
|
| 45 |
+
"""Vision tower, lm_head, the first/last `edge_layers` decoder layers and optionally all attention stay BF16."""
|
| 46 |
+
edges = [*range(edge_layers), *range(NUM_LAYERS - edge_layers, NUM_LAYERS)]
|
| 47 |
+
patterns = ["model.visual.*", "lm_head", *(f"model.language_model.layers.{i}.*" for i in edges)]
|
| 48 |
+
return patterns + (["model.language_model.layers.*.self_attn.*"] if skip_attention else [])
|
| 49 |
+
|
| 50 |
+
|
| 51 |
+
def forward_fn(pipe, device: str):
|
| 52 |
+
def forward(sample: dict):
|
| 53 |
+
kwargs = {
|
| 54 |
+
"prompt": sample["prompt"],
|
| 55 |
+
"num_inference_steps": 1, # the encoder runs once per call; one DiT step is enough
|
| 56 |
+
"output_type": "latent",
|
| 57 |
+
"generator": make_generator(sample["seed"], device=device),
|
| 58 |
+
}
|
| 59 |
+
if sample.get("image") is not None:
|
| 60 |
+
kwargs.update(image=sample["image"], output_resolution=sample["size"])
|
| 61 |
+
else:
|
| 62 |
+
kwargs.update(height=sample["size"], width=sample["size"])
|
| 63 |
+
return pipe(**kwargs)
|
| 64 |
+
|
| 65 |
+
return forward
|
| 66 |
+
|
| 67 |
+
|
| 68 |
+
def main():
|
| 69 |
+
parser = argparse.ArgumentParser()
|
| 70 |
+
parser.add_argument("--model-id", default="Qwen/Qwen-Image-2.1")
|
| 71 |
+
parser.add_argument("--edit-dataset", default="VyoJ/NHR-Edit-Change_Only")
|
| 72 |
+
parser.add_argument("--num-samples", type=int, default=128)
|
| 73 |
+
parser.add_argument("--rank", type=int, default=32)
|
| 74 |
+
parser.add_argument("--edge-layers", type=int, default=2, help="first/last decoder layers kept in BF16")
|
| 75 |
+
parser.add_argument("--skip-attention", action="store_true", help="keep all attention projections in BF16")
|
| 76 |
+
parser.add_argument("--no-gptq", action="store_true")
|
| 77 |
+
parser.add_argument("--out-dir", default="outputs/qwen21_nvfp4")
|
| 78 |
+
parser.add_argument("--cache-mode", choices=("reuse", "refresh", "disabled"), default="reuse")
|
| 79 |
+
parser.add_argument("--inspect", action="store_true")
|
| 80 |
+
args = parser.parse_args()
|
| 81 |
+
|
| 82 |
+
device = "cuda"
|
| 83 |
+
pipe = QwenImage21Pipeline.from_pretrained(args.model_id, dtype=torch.bfloat16).to(device)
|
| 84 |
+
pipe.set_progress_bar_config(disable=True)
|
| 85 |
+
encoder = pipe.text_encoder
|
| 86 |
+
scan = scan_linear_targets(
|
| 87 |
+
encoder, precision="nvfp4", rank=args.rank, skip=skip_patterns(args.edge_layers, args.skip_attention)
|
| 88 |
+
)
|
| 89 |
+
print(scan.format_text().splitlines()[0])
|
| 90 |
+
if args.inspect:
|
| 91 |
+
print(scan.format_text())
|
| 92 |
+
return
|
| 93 |
+
|
| 94 |
+
variant = f"r{args.rank}_e{args.edge_layers}{'_mlp' if args.skip_attention else ''}{'' if args.no_gptq else '-gptq'}"
|
| 95 |
+
tag = f"svdq-nvfp4_{variant}-qwen-image-2.1-text-encoder"
|
| 96 |
+
out = Path(args.out_dir)
|
| 97 |
+
cache = out / "calibration" / tag
|
| 98 |
+
spec = replace(svdquant_spec("nvfp4", compute_device=device), rank=args.rank, gptq=GptqSpec(enabled=not args.no_gptq))
|
| 99 |
+
quantize_and_export(
|
| 100 |
+
encoder,
|
| 101 |
+
spec,
|
| 102 |
+
scan.target_config,
|
| 103 |
+
CalibrationSpec(
|
| 104 |
+
samples=records(args.num_samples, args.edit_dataset),
|
| 105 |
+
num_samples=args.num_samples,
|
| 106 |
+
cache_num_samples=args.num_samples,
|
| 107 |
+
batch_size=1,
|
| 108 |
+
cache_dir=cache / "inputs",
|
| 109 |
+
cache_mode=args.cache_mode,
|
| 110 |
+
forward_fn=forward_fn(pipe, device),
|
| 111 |
+
max_rows_per_target=4096,
|
| 112 |
+
artifact_cache=QuantizationCacheSpec(cache / "artifacts", args.cache_mode),
|
| 113 |
+
scope_capture_mode="all_targets",
|
| 114 |
+
),
|
| 115 |
+
ExportSpec(output=out / "checkpoints" / f"{tag}.safetensors"),
|
| 116 |
+
LoggingConfig(name=tag),
|
| 117 |
+
)
|
| 118 |
+
|
| 119 |
+
|
| 120 |
+
if __name__ == "__main__":
|
| 121 |
+
main()
|
transformer/config.json
CHANGED
|
@@ -1,7 +1,6 @@
|
|
| 1 |
{
|
| 2 |
"_class_name": "QwenImage21Transformer2DModel",
|
| 3 |
-
"_diffusers_version": "0.
|
| 4 |
-
"_name_or_path": "/mnt/hf_cache/hub/models--Qwen--Qwen-Image-2.1/snapshots/b3179ad355be050328e483a9dfdd9e60cd62adfa/transformer",
|
| 5 |
"attention_head_dim": 128,
|
| 6 |
"axes_dims_rope": [
|
| 7 |
16,
|
|
@@ -18,43 +17,210 @@
|
|
| 18 |
"out_channels": 64,
|
| 19 |
"patch_size": 1,
|
| 20 |
"quantization_config": {
|
| 21 |
-
"
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
"
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
"
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 59 |
}
|
| 60 |
-
}
|
|
|
|
| 1 |
{
|
| 2 |
"_class_name": "QwenImage21Transformer2DModel",
|
| 3 |
+
"_diffusers_version": "0.37.0.dev0",
|
|
|
|
| 4 |
"attention_head_dim": 128,
|
| 5 |
"axes_dims_rope": [
|
| 6 |
16,
|
|
|
|
| 17 |
"out_channels": 64,
|
| 18 |
"patch_size": 1,
|
| 19 |
"quantization_config": {
|
| 20 |
+
"compute_dtype": "bfloat16",
|
| 21 |
+
"quant_method": "nunchaku_lite",
|
| 22 |
+
"svdq_w4a4": {
|
| 23 |
+
"group_size": 16,
|
| 24 |
+
"precision": "nvfp4",
|
| 25 |
+
"rank": 32,
|
| 26 |
+
"targets": [
|
| 27 |
+
"transformer_blocks.2.attn.to_q",
|
| 28 |
+
"transformer_blocks.2.attn.to_k",
|
| 29 |
+
"transformer_blocks.2.attn.to_v",
|
| 30 |
+
"transformer_blocks.2.attn.to_out.0",
|
| 31 |
+
"transformer_blocks.2.img_mlp.proj",
|
| 32 |
+
"transformer_blocks.2.img_mlp.out",
|
| 33 |
+
"transformer_blocks.2.img_mlp.gate_layer",
|
| 34 |
+
"transformer_blocks.3.attn.to_q",
|
| 35 |
+
"transformer_blocks.3.attn.to_k",
|
| 36 |
+
"transformer_blocks.3.attn.to_v",
|
| 37 |
+
"transformer_blocks.3.attn.to_out.0",
|
| 38 |
+
"transformer_blocks.3.img_mlp.proj",
|
| 39 |
+
"transformer_blocks.3.img_mlp.out",
|
| 40 |
+
"transformer_blocks.3.img_mlp.gate_layer",
|
| 41 |
+
"transformer_blocks.4.attn.to_q",
|
| 42 |
+
"transformer_blocks.4.attn.to_k",
|
| 43 |
+
"transformer_blocks.4.attn.to_v",
|
| 44 |
+
"transformer_blocks.4.attn.to_out.0",
|
| 45 |
+
"transformer_blocks.4.img_mlp.proj",
|
| 46 |
+
"transformer_blocks.4.img_mlp.out",
|
| 47 |
+
"transformer_blocks.4.img_mlp.gate_layer",
|
| 48 |
+
"transformer_blocks.5.attn.to_q",
|
| 49 |
+
"transformer_blocks.5.attn.to_k",
|
| 50 |
+
"transformer_blocks.5.attn.to_v",
|
| 51 |
+
"transformer_blocks.5.attn.to_out.0",
|
| 52 |
+
"transformer_blocks.5.img_mlp.proj",
|
| 53 |
+
"transformer_blocks.5.img_mlp.out",
|
| 54 |
+
"transformer_blocks.5.img_mlp.gate_layer",
|
| 55 |
+
"transformer_blocks.6.attn.to_q",
|
| 56 |
+
"transformer_blocks.6.attn.to_k",
|
| 57 |
+
"transformer_blocks.6.attn.to_v",
|
| 58 |
+
"transformer_blocks.6.attn.to_out.0",
|
| 59 |
+
"transformer_blocks.6.img_mlp.proj",
|
| 60 |
+
"transformer_blocks.6.img_mlp.out",
|
| 61 |
+
"transformer_blocks.6.img_mlp.gate_layer",
|
| 62 |
+
"transformer_blocks.7.attn.to_q",
|
| 63 |
+
"transformer_blocks.7.attn.to_k",
|
| 64 |
+
"transformer_blocks.7.attn.to_v",
|
| 65 |
+
"transformer_blocks.7.attn.to_out.0",
|
| 66 |
+
"transformer_blocks.7.img_mlp.proj",
|
| 67 |
+
"transformer_blocks.7.img_mlp.out",
|
| 68 |
+
"transformer_blocks.7.img_mlp.gate_layer",
|
| 69 |
+
"transformer_blocks.8.attn.to_q",
|
| 70 |
+
"transformer_blocks.8.attn.to_k",
|
| 71 |
+
"transformer_blocks.8.attn.to_v",
|
| 72 |
+
"transformer_blocks.8.attn.to_out.0",
|
| 73 |
+
"transformer_blocks.8.img_mlp.proj",
|
| 74 |
+
"transformer_blocks.8.img_mlp.out",
|
| 75 |
+
"transformer_blocks.8.img_mlp.gate_layer",
|
| 76 |
+
"transformer_blocks.9.attn.to_q",
|
| 77 |
+
"transformer_blocks.9.attn.to_k",
|
| 78 |
+
"transformer_blocks.9.attn.to_v",
|
| 79 |
+
"transformer_blocks.9.attn.to_out.0",
|
| 80 |
+
"transformer_blocks.9.img_mlp.proj",
|
| 81 |
+
"transformer_blocks.9.img_mlp.out",
|
| 82 |
+
"transformer_blocks.9.img_mlp.gate_layer",
|
| 83 |
+
"transformer_blocks.10.attn.to_q",
|
| 84 |
+
"transformer_blocks.10.attn.to_k",
|
| 85 |
+
"transformer_blocks.10.attn.to_v",
|
| 86 |
+
"transformer_blocks.10.attn.to_out.0",
|
| 87 |
+
"transformer_blocks.10.img_mlp.proj",
|
| 88 |
+
"transformer_blocks.10.img_mlp.out",
|
| 89 |
+
"transformer_blocks.10.img_mlp.gate_layer",
|
| 90 |
+
"transformer_blocks.11.attn.to_q",
|
| 91 |
+
"transformer_blocks.11.attn.to_k",
|
| 92 |
+
"transformer_blocks.11.attn.to_v",
|
| 93 |
+
"transformer_blocks.11.attn.to_out.0",
|
| 94 |
+
"transformer_blocks.11.img_mlp.proj",
|
| 95 |
+
"transformer_blocks.11.img_mlp.out",
|
| 96 |
+
"transformer_blocks.11.img_mlp.gate_layer",
|
| 97 |
+
"transformer_blocks.12.attn.to_q",
|
| 98 |
+
"transformer_blocks.12.attn.to_k",
|
| 99 |
+
"transformer_blocks.12.attn.to_v",
|
| 100 |
+
"transformer_blocks.12.attn.to_out.0",
|
| 101 |
+
"transformer_blocks.12.img_mlp.proj",
|
| 102 |
+
"transformer_blocks.12.img_mlp.out",
|
| 103 |
+
"transformer_blocks.12.img_mlp.gate_layer",
|
| 104 |
+
"transformer_blocks.13.attn.to_q",
|
| 105 |
+
"transformer_blocks.13.attn.to_k",
|
| 106 |
+
"transformer_blocks.13.attn.to_v",
|
| 107 |
+
"transformer_blocks.13.attn.to_out.0",
|
| 108 |
+
"transformer_blocks.13.img_mlp.proj",
|
| 109 |
+
"transformer_blocks.13.img_mlp.out",
|
| 110 |
+
"transformer_blocks.13.img_mlp.gate_layer",
|
| 111 |
+
"transformer_blocks.14.attn.to_q",
|
| 112 |
+
"transformer_blocks.14.attn.to_k",
|
| 113 |
+
"transformer_blocks.14.attn.to_v",
|
| 114 |
+
"transformer_blocks.14.attn.to_out.0",
|
| 115 |
+
"transformer_blocks.14.img_mlp.proj",
|
| 116 |
+
"transformer_blocks.14.img_mlp.out",
|
| 117 |
+
"transformer_blocks.14.img_mlp.gate_layer",
|
| 118 |
+
"transformer_blocks.15.attn.to_q",
|
| 119 |
+
"transformer_blocks.15.attn.to_k",
|
| 120 |
+
"transformer_blocks.15.attn.to_v",
|
| 121 |
+
"transformer_blocks.15.attn.to_out.0",
|
| 122 |
+
"transformer_blocks.15.img_mlp.proj",
|
| 123 |
+
"transformer_blocks.15.img_mlp.out",
|
| 124 |
+
"transformer_blocks.15.img_mlp.gate_layer",
|
| 125 |
+
"transformer_blocks.16.attn.to_q",
|
| 126 |
+
"transformer_blocks.16.attn.to_k",
|
| 127 |
+
"transformer_blocks.16.attn.to_v",
|
| 128 |
+
"transformer_blocks.16.attn.to_out.0",
|
| 129 |
+
"transformer_blocks.16.img_mlp.proj",
|
| 130 |
+
"transformer_blocks.16.img_mlp.out",
|
| 131 |
+
"transformer_blocks.16.img_mlp.gate_layer",
|
| 132 |
+
"transformer_blocks.17.attn.to_q",
|
| 133 |
+
"transformer_blocks.17.attn.to_k",
|
| 134 |
+
"transformer_blocks.17.attn.to_v",
|
| 135 |
+
"transformer_blocks.17.attn.to_out.0",
|
| 136 |
+
"transformer_blocks.17.img_mlp.proj",
|
| 137 |
+
"transformer_blocks.17.img_mlp.out",
|
| 138 |
+
"transformer_blocks.17.img_mlp.gate_layer",
|
| 139 |
+
"transformer_blocks.18.attn.to_q",
|
| 140 |
+
"transformer_blocks.18.attn.to_k",
|
| 141 |
+
"transformer_blocks.18.attn.to_v",
|
| 142 |
+
"transformer_blocks.18.attn.to_out.0",
|
| 143 |
+
"transformer_blocks.18.img_mlp.proj",
|
| 144 |
+
"transformer_blocks.18.img_mlp.out",
|
| 145 |
+
"transformer_blocks.18.img_mlp.gate_layer",
|
| 146 |
+
"transformer_blocks.19.attn.to_q",
|
| 147 |
+
"transformer_blocks.19.attn.to_k",
|
| 148 |
+
"transformer_blocks.19.attn.to_v",
|
| 149 |
+
"transformer_blocks.19.attn.to_out.0",
|
| 150 |
+
"transformer_blocks.19.img_mlp.proj",
|
| 151 |
+
"transformer_blocks.19.img_mlp.out",
|
| 152 |
+
"transformer_blocks.19.img_mlp.gate_layer",
|
| 153 |
+
"transformer_blocks.20.attn.to_q",
|
| 154 |
+
"transformer_blocks.20.attn.to_k",
|
| 155 |
+
"transformer_blocks.20.attn.to_v",
|
| 156 |
+
"transformer_blocks.20.attn.to_out.0",
|
| 157 |
+
"transformer_blocks.20.img_mlp.proj",
|
| 158 |
+
"transformer_blocks.20.img_mlp.out",
|
| 159 |
+
"transformer_blocks.20.img_mlp.gate_layer",
|
| 160 |
+
"transformer_blocks.21.attn.to_q",
|
| 161 |
+
"transformer_blocks.21.attn.to_k",
|
| 162 |
+
"transformer_blocks.21.attn.to_v",
|
| 163 |
+
"transformer_blocks.21.attn.to_out.0",
|
| 164 |
+
"transformer_blocks.21.img_mlp.proj",
|
| 165 |
+
"transformer_blocks.21.img_mlp.out",
|
| 166 |
+
"transformer_blocks.21.img_mlp.gate_layer",
|
| 167 |
+
"transformer_blocks.22.attn.to_q",
|
| 168 |
+
"transformer_blocks.22.attn.to_k",
|
| 169 |
+
"transformer_blocks.22.attn.to_v",
|
| 170 |
+
"transformer_blocks.22.attn.to_out.0",
|
| 171 |
+
"transformer_blocks.22.img_mlp.proj",
|
| 172 |
+
"transformer_blocks.22.img_mlp.out",
|
| 173 |
+
"transformer_blocks.22.img_mlp.gate_layer",
|
| 174 |
+
"transformer_blocks.23.attn.to_q",
|
| 175 |
+
"transformer_blocks.23.attn.to_k",
|
| 176 |
+
"transformer_blocks.23.attn.to_v",
|
| 177 |
+
"transformer_blocks.23.attn.to_out.0",
|
| 178 |
+
"transformer_blocks.23.img_mlp.proj",
|
| 179 |
+
"transformer_blocks.23.img_mlp.out",
|
| 180 |
+
"transformer_blocks.23.img_mlp.gate_layer",
|
| 181 |
+
"transformer_blocks.24.attn.to_q",
|
| 182 |
+
"transformer_blocks.24.attn.to_k",
|
| 183 |
+
"transformer_blocks.24.attn.to_v",
|
| 184 |
+
"transformer_blocks.24.attn.to_out.0",
|
| 185 |
+
"transformer_blocks.24.img_mlp.proj",
|
| 186 |
+
"transformer_blocks.24.img_mlp.out",
|
| 187 |
+
"transformer_blocks.24.img_mlp.gate_layer",
|
| 188 |
+
"transformer_blocks.25.attn.to_q",
|
| 189 |
+
"transformer_blocks.25.attn.to_k",
|
| 190 |
+
"transformer_blocks.25.attn.to_v",
|
| 191 |
+
"transformer_blocks.25.attn.to_out.0",
|
| 192 |
+
"transformer_blocks.25.img_mlp.proj",
|
| 193 |
+
"transformer_blocks.25.img_mlp.out",
|
| 194 |
+
"transformer_blocks.25.img_mlp.gate_layer",
|
| 195 |
+
"transformer_blocks.26.attn.to_q",
|
| 196 |
+
"transformer_blocks.26.attn.to_k",
|
| 197 |
+
"transformer_blocks.26.attn.to_v",
|
| 198 |
+
"transformer_blocks.26.attn.to_out.0",
|
| 199 |
+
"transformer_blocks.26.img_mlp.proj",
|
| 200 |
+
"transformer_blocks.26.img_mlp.out",
|
| 201 |
+
"transformer_blocks.26.img_mlp.gate_layer",
|
| 202 |
+
"transformer_blocks.27.attn.to_q",
|
| 203 |
+
"transformer_blocks.27.attn.to_k",
|
| 204 |
+
"transformer_blocks.27.attn.to_v",
|
| 205 |
+
"transformer_blocks.27.attn.to_out.0",
|
| 206 |
+
"transformer_blocks.27.img_mlp.proj",
|
| 207 |
+
"transformer_blocks.27.img_mlp.out",
|
| 208 |
+
"transformer_blocks.27.img_mlp.gate_layer",
|
| 209 |
+
"transformer_blocks.28.attn.to_q",
|
| 210 |
+
"transformer_blocks.28.attn.to_k",
|
| 211 |
+
"transformer_blocks.28.attn.to_v",
|
| 212 |
+
"transformer_blocks.28.attn.to_out.0",
|
| 213 |
+
"transformer_blocks.28.img_mlp.proj",
|
| 214 |
+
"transformer_blocks.28.img_mlp.out",
|
| 215 |
+
"transformer_blocks.28.img_mlp.gate_layer",
|
| 216 |
+
"transformer_blocks.29.attn.to_q",
|
| 217 |
+
"transformer_blocks.29.attn.to_k",
|
| 218 |
+
"transformer_blocks.29.attn.to_v",
|
| 219 |
+
"transformer_blocks.29.attn.to_out.0",
|
| 220 |
+
"transformer_blocks.29.img_mlp.proj",
|
| 221 |
+
"transformer_blocks.29.img_mlp.out",
|
| 222 |
+
"transformer_blocks.29.img_mlp.gate_layer"
|
| 223 |
+
]
|
| 224 |
+
}
|
| 225 |
}
|
| 226 |
+
}
|
transformer/diffusion_pytorch_model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:0304fc3e84b13663b439601380066754ef9b92025e159275abe163bc5f1ceeb4
|
| 3 |
+
size 5603299808
|
transformer/modeling_nunchaku_qwenimage21.py
ADDED
|
@@ -0,0 +1,14 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Qwen-Image-2.1 transformer, unchanged, loaded as a custom component.
|
| 2 |
+
|
| 3 |
+
Its only job is importing `nunchaku_kernels`, so the NVFP4 kernels are set up even when
|
| 4 |
+
Diffusers loads this component before the text encoder. The Nunchaku Lite quantization
|
| 5 |
+
itself is handled by Diffusers' built-in quantizer from `config.json`.
|
| 6 |
+
"""
|
| 7 |
+
|
| 8 |
+
from diffusers import QwenImage21Transformer2DModel
|
| 9 |
+
|
| 10 |
+
from .nunchaku_kernels import KERNELS # noqa: F401
|
| 11 |
+
|
| 12 |
+
|
| 13 |
+
class NunchakuQwenImage21Transformer2DModel(QwenImage21Transformer2DModel):
|
| 14 |
+
pass
|
transformer/nunchaku_kernels.py
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Shared by text_encoder/ and transformer/: whichever component Diffusers loads first sets this up.
|
| 2 |
+
#
|
| 3 |
+
# Diffusers fetches the Nunchaku Lite kernels from `rootonchair/nunchaku-lite-kernels`, which is
|
| 4 |
+
# no longer downloadable. Loading this repo with trust_remote_code=True already runs this code,
|
| 5 |
+
# so point that kernel name at the rebuilt copy before Diffusers imports its Nunchaku utilities.
|
| 6 |
+
import os
|
| 7 |
+
|
| 8 |
+
from huggingface_hub import snapshot_download
|
| 9 |
+
|
| 10 |
+
KERNELS = "rootonchair/nunchaku-lite-kernels"
|
| 11 |
+
if KERNELS not in os.environ.get("LOCAL_KERNELS", ""):
|
| 12 |
+
local = f"{KERNELS}={snapshot_download('joseplcam/nunchaku-lite-kernels')}"
|
| 13 |
+
os.environ["LOCAL_KERNELS"] = ":".join(filter(None, [os.environ.get("LOCAL_KERNELS"), local]))
|
| 14 |
+
if "DIFFUSERS_TRUST_REMOTE_KERNELS" not in os.environ:
|
| 15 |
+
# Diffusers read this variable into a constant at import time, so set both.
|
| 16 |
+
import diffusers.utils.constants
|
| 17 |
+
|
| 18 |
+
os.environ["DIFFUSERS_TRUST_REMOTE_KERNELS"] = "true"
|
| 19 |
+
diffusers.utils.constants.DIFFUSERS_TRUST_REMOTE_KERNELS = True
|
vae/config.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
{
|
| 2 |
"_class_name": "AutoencoderKLQwenImage21",
|
| 3 |
-
"_diffusers_version": "0.
|
| 4 |
"attn_scales": [],
|
| 5 |
"base_dim": 96,
|
| 6 |
"decoder_base_dim": 144,
|
|
|
|
| 1 |
{
|
| 2 |
"_class_name": "AutoencoderKLQwenImage21",
|
| 3 |
+
"_diffusers_version": "0.41.0.dev0",
|
| 4 |
"attn_scales": [],
|
| 5 |
"base_dim": 96,
|
| 6 |
"decoder_base_dim": 144,
|
vae/diffusion_pytorch_model.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:71879ffd5321e6d10c3c87513e2b474b1252efa7f3dec2969214a9bf06a6dd5c
|
| 3 |
+
size 675508656
|