joseplcam commited on
Commit
7cbf234
·
verified ·
1 Parent(s): 3814de5

Replace with SVDQuant+GPTQ NVFP4 DiT and text encoder, native in Diffusers

Browse files
NOTICE CHANGED
@@ -2,22 +2,32 @@ Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 H
2
 
3
  Built with Qwen.
4
 
5
- This repository redistributes, for non-commercial research and evaluation use only, files derived from:
 
6
 
7
- - Qwen/Qwen-Image-2.1 (revision 790c92633540aa0cb11d9abf19eb46d861714758)
8
- model_index.json, LICENSE, processor/, scheduler/, vae/, text_encoder/config.json,
9
- text_encoder/generation_config.json. Unmodified.
10
 
11
- - HangGlidersRule/Darkstar-Qwen-Image-2.1-Base-ModelOpt-W4A4-NVFP4 (revision 3b7aa842dd9c1ec8f90d005979f8b2a0d7d7e95b)
12
- transformer/config.json, transformer/diffusion_pytorch_model.safetensors. Unmodified
13
- (sha256 a6c3080b764344729ef95846faa3a0a01eb36bdef93f785ed07f4ac5751ad76a).
14
 
15
- - BennyDaBall/Qwen-Image-2.1-NVFP4 (revision 1a38d44a3a2f35cb0b543a25b04da0a963e7b5e6)
16
- text_encoders/qwen3vl_8b_nvfp4.safetensors
17
- (sha256 cdd9b85bcb60d5cb358ddecfd45a48fa15cb728b70d88bd2028840eaef095d16).
 
 
18
 
19
- MODIFIED FILE: text_encoder/model.safetensors was changed by joseplcam from the BennyDaBall file above.
20
- The only change is tensor names: the prefixes "model.layers.", "model.embed_tokens." and "model.norm."
21
- were renamed to "model.language_model.layers.", "model.language_model.embed_tokens." and
22
- "model.language_model.norm." to match the Hugging Face Qwen3-VL layout. Tensor data, dtypes, shapes and
23
- file metadata are unchanged (sha256 of the result 42882503fe37fbc72f374fb9e55e79f020f9cecca02463417176fdd902be470e).
 
 
 
 
 
 
 
 
 
 
 
 
2
 
3
  Built with Qwen.
4
 
5
+ This repository redistributes, for non-commercial research and evaluation use only, files derived from
6
+ Qwen/Qwen-Image-2.1 (revision 790c92633540aa0cb11d9abf19eb46d861714758).
7
 
8
+ Unmodified files: LICENSE, processor/, scheduler/, text_encoder/generation_config.json.
 
 
9
 
10
+ MODIFIED FILES (changed by joseplcam):
 
 
11
 
12
+ - transformer/diffusion_pytorch_model.safetensors, transformer/config.json
13
+ Quantized from the official BF16 transformer with SVDQuant + GPTQ (NVFP4 weights and activations,
14
+ rank-32 BF16 low-rank branch) using diffuse-compressor. Blocks 0, 1, 30, 31 and the global
15
+ modulation are the official BF16 tensors. config.json gained a `quantization_config` entry.
16
+ sha256 0304fc3e84b13663b439601380066754ef9b92025e159275abe163bc5f1ceeb4
17
 
18
+ - text_encoder/model.safetensors, text_encoder/config.json
19
+ Quantized from the official BF16 Qwen3-VL text encoder with the same method: the MLP projections
20
+ (gate/up/down) of decoder layers 4-31, rank-128 low-rank branch. All attention projections, decoder
21
+ layers 0-3 and 32-35, the vision tower, embeddings and lm_head are the official BF16 tensors.
22
+ config.json gained a `nunchaku_lite` entry.
23
+ sha256 83c4462b5ce6b98e2708ca53031e2f9a60a8e402c4c4a0ccd19cb2df8e817ae2
24
+
25
+ - vae/diffusion_pytorch_model.safetensors, vae/config.json
26
+ The official VAE weights cast from FP32 to BF16.
27
+ sha256 71879ffd5321e6d10c3c87513e2b474b1252efa7f3dec2969214a9bf06a6dd5c
28
+
29
+ - model_index.json
30
+ The `text_encoder` and `transformer` entries point to the custom classes below.
31
+
32
+ NEW FILES: text_encoder/modeling_nunchaku_qwen3vl.py, transformer/modeling_nunchaku_qwenimage21.py,
33
+ text_encoder/nunchaku_kernels.py, transformer/nunchaku_kernels.py.
README.md CHANGED
@@ -10,129 +10,150 @@ tags:
10
  - qwen-image
11
  - qwen-image-2.1
12
  - nvfp4
13
- - sglang
 
14
  - blackwell
15
  - text-to-image
16
  - image-editing
17
  ---
18
 
19
- # Qwen-Image-2.1-NVFP4 (DiT + text encoder) for SGLang Diffusion
20
 
21
  **Built with Qwen.** Non-commercial research and evaluation use only (Qwen Research License, see `LICENSE` and `NOTICE`).
22
 
23
- This repository is a ready-to-run combination of existing NVFP4 quantizations of
24
- [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1), packaged in the diffusers layout so that
25
- [SGLang Diffusion](https://github.com/sgl-project/sglang) can load it from one model id and run both large
26
- components on native Blackwell FP4 kernels. No new quantization was done here: the value of this repo is the tested
27
- combination, the small key-rename that makes the encoder loadable by SGLang, and the measurements below.
28
 
29
- Text-to-image, image editing and RGBA output all work as in the original pipeline.
 
 
30
 
31
- ## What is inside, and where each part comes from
32
-
33
- | Path | Precision | Source | Changes |
34
- |---|---|---|---|
35
- | `transformer/` (7B single-stream DiT) | NVFP4 W4A4, group 16, FP8 block scales. Blocks 0, 1, 30, 31 and the embeddings, modulation and output layers stay BF16 | [HangGlidersRule/Darkstar-Qwen-Image-2.1-Base-ModelOpt-W4A4-NVFP4](https://huggingface.co/HangGlidersRule/Darkstar-Qwen-Image-2.1-Base-ModelOpt-W4A4-NVFP4) @ `3b7aa84`. NVIDIA ModelOpt export of the base model | None (byte-identical) |
36
- | `text_encoder/model.safetensors` (Qwen3-VL 8B) | NVFP4 for the 252 language-model linear layers; vision tower, embeddings and `lm_head` stay BF16 | [BennyDaBall/Qwen-Image-2.1-NVFP4](https://huggingface.co/BennyDaBall/Qwen-Image-2.1-NVFP4) @ `1a38d44`, file `text_encoders/qwen3vl_8b_nvfp4.safetensors`. ComfyUI-style NVFP4 converted from the official BF16 weights | Tensor names renamed to the Hugging Face Qwen3-VL layout (`model.layers.*` → `model.language_model.layers.*`, same for `embed_tokens` and `norm`). Data unchanged |
37
- | `vae/`, `processor/`, `scheduler/`, `model_index.json`, `text_encoder/*.json`, `LICENSE` | BF16 / FP32 as released | [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) @ `790c926` | None |
38
-
39
- Why these two parts:
40
-
41
- - The DiT is a ModelOpt NVFP4 export. SGLang detects it from `transformer/config.json` and runs it with FlashInfer's
42
- CUTLASS FP4 GEMM, which has an SM120 path (RTX 50 series, RTX PRO 6000). Keeping the first two and last two blocks in
43
- BF16 preserves quality: its author reports parity with BF16 on a 25-case evaluation.
44
- - The encoder is the only published NVFP4 Qwen3-VL 8B built from the official weights (not an abliterated variant).
45
- SGLang runs its Comfy NVFP4 layers natively via `comfy-kitchen` FP4 matmuls. The rename was needed because SGLang's
46
- native Qwen3-VL encoder expects the Hugging Face tensor names.
47
-
48
- Total size is about 14 GB, against about 32 GB for the BF16 release.
49
-
50
- ## Usage (SGLang Diffusion)
51
-
52
- Tested with SGLang commit `ffac53d779c08dcdab2d07e5e2a41dba83f0e65c`, PyTorch 2.13.0+cu130, FlashInfer 0.6.18 and
53
- `comfy-kitchen==0.2.35`. You need a Blackwell GPU (compute capability 10.0 or 12.0).
54
-
55
- SGLang only recognizes the Comfy NVFP4 encoder when its weights path points at the file, so pass it explicitly:
56
 
57
  ```python
58
- from sglang.multimodal_gen import DiffGenerator
59
-
60
- repo = "joseplcam/Qwen-Image-2.1-NVFP4"
61
- with DiffGenerator.from_pretrained(
62
- model_path=repo,
63
- component_weights_paths={"text_encoder": f"{repo}/text_encoder/model.safetensors"},
64
- performance_mode="speed", # keep every component resident on the GPU
65
- attention_backend="torch_sdpa", # "sage_attn" is faster at 1024x1024 and above
66
- ) as generator:
67
- result = generator.generate(sampling_params_kwargs=dict(
68
- prompt="A capybara reading a book by candlelight",
69
- width=512, height=512, num_inference_steps=40, guidance_scale=1.0, seed=42,
70
- ))
71
- ```
72
-
73
- Add `image_path="input.png"` for editing. The same arguments work with a local clone of this repository.
74
 
75
- Serving with batching of concurrent text-to-image requests:
 
 
76
 
77
- ```bash
78
- sglang serve --model-path joseplcam/Qwen-Image-2.1-NVFP4 \
79
- --component-weights-paths.text_encoder joseplcam/Qwen-Image-2.1-NVFP4/text_encoder/model.safetensors \
80
- --performance-mode speed --batching-max-size 8 --batching-delay-ms 50
81
  ```
82
 
83
- Tips:
84
-
85
- - Set `performance_mode="speed"`. Without it, SGLang's automatic memory policy may stream the text encoder layer by
86
- layer from host memory even on a 96 GB GPU, which made every request about 10x slower in our tests.
87
- - The first run JIT-compiles FlashInfer's FP4 kernels (about 5 minutes; cached in `~/.cache/sglang`). On machines with
88
- about 32 GB of RAM, set `MAX_JOBS=4`, or the compiler may run out of memory.
89
- - On NixOS, the JIT linker needs `libcuda` and `libcudart`, for example
90
- `LIBRARY_PATH=/run/opengl-driver/lib:<cuda>/lib`.
91
- - torch.compile made this model slower in every configuration we tried, so leave it off.
92
-
93
- ## Measurements (RTX PRO 6000 Blackwell 96 GB, SGLang serve, 512x512, 40 steps)
94
-
95
- Average of 3 runs after a warmup. Batch size is the number of concurrent requests, each with a different prompt
96
- (and a different reference image for edits). VRAM is the peak for the whole GPU.
97
-
98
- Text-to-image:
99
-
100
- | Batch | SDPA | SageAttention2 | SDPA + torch.compile | SageAttention2 + torch.compile |
101
- |---|---|---|---|---|
102
- | 1 | **2.42 s** (0.41 img/s, 16.7 GB) | 2.54 s (0.39 img/s, 16.7 GB) | 6.39 s (0.16 img/s) | 6.80 s (0.15 img/s) |
103
- | 2 | **2.86 s** (0.70 img/s, 17.9 GB) | 3.02 s (0.66 img/s, 17.9 GB) | 5.75 s (0.35 img/s) | 6.17 s (0.32 img/s) |
104
- | 4 | **5.58 s** (0.72 img/s, 20.1 GB) | 5.82 s (0.69 img/s, 20.1 GB) | 6.39 s (0.63 img/s) | 7.77 s (0.52 img/s) |
105
- | 8 | **11.33 s** (0.71 img/s, 25.0 GB) | 11.70 s (0.68 img/s, 25.1 GB) | 12.34 s (0.65 img/s) | 14.35 s (0.56 img/s) |
106
-
107
- Image editing runs at about 2.4 s per image (0.41 img/s, 16.8 GB) at any batch size, because SGLang does not merge
108
- edit requests that carry reference images; they run one after another.
109
-
110
- Comparison with other precisions (single request, SDPA):
111
-
112
- | Setup | 512x512 | 1024x1024 | Peak VRAM at 512x512 |
113
- |---|---|---|---|
114
- | BF16 DiT + BF16 encoder (official release) | 2.85 s | 11.45 s | 34.7 GB |
115
- | NVFP4 DiT + BF16 encoder | 2.26 s | 6.00 s | 26.7 GB |
116
- | NVFP4 DiT + NVFP4 encoder (this repo) | 2.29 s | not measured | 16.9 GB |
117
-
118
- At 1024x1024 with the NVFP4 DiT, SageAttention2 took 5.53 s against 6.00 s for SDPA.
119
-
120
- Quality: we saw no artifacts on our sample prompts or edits. Quantization changes composition details compared with
121
- BF16 at the same seed (LPIPS against BF16: 0.26 at 512x512 and 0.11 at 1024x1024 for the NVFP4 DiT; 0.19 for adding
122
- the NVFP4 encoder). Check quality on your own prompts, especially for text rendering.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
123
 
124
  ## Limitations
125
 
 
 
 
126
  - Non-commercial research and evaluation use only.
127
- - Measured on one GPU model only (RTX PRO 6000 Blackwell). Other Blackwell GPUs are untested.
128
- - The encoder needs the explicit `component_weights_paths` argument shown above.
129
- - diffusers `from_pretrained` does not load the NVFP4 components; use SGLang.
 
 
 
 
130
 
131
  ## Credits and license
132
 
133
- - Qwen team: Qwen-Image 2.1 model, VAE, processor and configuration.
134
- - HangGlidersRule (Model Forge): ModelOpt NVFP4 DiT.
135
- - BennyDaBall: NVFP4 Qwen3-VL 8B text encoder.
 
136
 
137
- Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology
138
- Co., Ltd. All Rights Reserved. See `LICENSE` and `NOTICE`, which also lists the one modified file.
 
10
  - qwen-image
11
  - qwen-image-2.1
12
  - nvfp4
13
+ - svdquant
14
+ - nunchaku
15
  - blackwell
16
  - text-to-image
17
  - image-editing
18
  ---
19
 
20
+ # Qwen-Image-2.1-NVFP4 (SVDQuant, DiT + text encoder, native in Diffusers)
21
 
22
  **Built with Qwen.** Non-commercial research and evaluation use only (Qwen Research License, see `LICENSE` and `NOTICE`).
23
 
24
+ An NVFP4 quantization of [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1), made
25
+ from the official BF16 weights with SVDQuant + GPTQ and calibrated on both text-to-image prompts and
26
+ image edits. Both large components run on native Blackwell FP4 tensor cores, and the pipeline loads with a
27
+ single `DiffusionPipeline.from_pretrained` call.
 
28
 
29
+ At 512x512 on an RTX PRO 6000 Blackwell it is about **1.7x faster than BF16 and needs about 40% less VRAM**,
30
+ and it matches BF16 closely on text-to-image, typography and edits, including a hard recolor edit that
31
+ earlier 4-bit text encoders failed.
32
 
33
+ ## Usage
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
34
 
35
  ```python
36
+ import torch
37
+ from diffusers import DiffusionPipeline
 
 
 
 
 
 
 
 
 
 
 
 
 
 
38
 
39
+ pipe = DiffusionPipeline.from_pretrained(
40
+ "joseplcam/Qwen-Image-2.1-NVFP4", dtype=torch.bfloat16, trust_remote_code=True
41
+ ).to("cuda")
42
 
43
+ image = pipe("A capybara reading a book by candlelight", height=512, width=512, num_inference_steps=40).images[0]
44
+ edited = pipe("Make it night time with moonlight", image=image, output_resolution=512, num_inference_steps=40).images[0]
 
 
45
  ```
46
 
47
+ `trust_remote_code=True` loads two small files from this repository:
48
+
49
+ - `text_encoder/modeling_nunchaku_qwen3vl.py`: a `Qwen3VLForConditionalGeneration` subclass that swaps the
50
+ quantized linears for Diffusers' own `SVDQW4A4Linear` before loading the weights.
51
+ - `nunchaku_kernels.py` (in `text_encoder/` and `transformer/`): Diffusers loads the NVFP4 kernels from
52
+ `rootonchair/nunchaku-lite-kernels`, which is no longer downloadable. This file points that name at
53
+ [joseplcam/nunchaku-lite-kernels](https://huggingface.co/joseplcam/nunchaku-lite-kernels), an unmodified
54
+ build of the same open-source kernels, and sets `DIFFUSERS_TRUST_REMOTE_KERNELS=true` unless you already set
55
+ it. Set `LOCAL_KERNELS` yourself to use a different build.
56
+
57
+ Use a different seed for an edit than the one that generated its input image. Qwen-Image-2.1 returns an
58
+ over-sharpened copy that ignores the prompt when the edit starts from the same noise
59
+ ([diffusers #14824](https://github.com/huggingface/diffusers/issues/14824)); this is a base-model behaviour.
60
+
61
+ ### Requirements
62
+
63
+ - NVIDIA Blackwell GPU with compute capability 12.0 (RTX 50 series, RTX PRO 6000). The published kernel build
64
+ targets `sm_120a` only.
65
+ - Linux x86_64, Python 3.12, PyTorch 2.13 with CUDA 13.0, which the kernel build targets.
66
+ - Diffusers from `main` with Qwen-Image-2.1 and Nunchaku Lite support (tested at commit `0377f0c`),
67
+ `transformers>=5.12`, `accelerate`, `kernels>=0.14`.
68
+
69
+ ## What is inside
70
+
71
+ | Component | Precision | Details |
72
+ |---|---|---|
73
+ | `transformer/` (7B DiT) | NVFP4 W4A4, group 16, FP8 block scales + BF16 rank-32 low-rank branch | All attention and MLP projections of blocks 2-29 (196 layers). Blocks 0, 1, 30, 31 and the global modulation stay BF16. |
74
+ | `text_encoder/` (Qwen3-VL 8B) | NVFP4 W4A4 + BF16 rank-128 low-rank branch | MLP projections (gate/up/down) of decoder layers 4-31 (84 layers). All attention projections, layers 0-3 and 32-35, the vision tower, embeddings and `lm_head` stay BF16. |
75
+ | `vae/` | BF16 | Official weights cast from FP32. |
76
+ | `processor/`, `scheduler/` | as released | Unmodified. |
77
+
78
+ Total download: about 17 GB, against about 32 GB for the BF16 release.
79
+
80
+ ## How it was made
81
+
82
+ - **Method:** SVDQuant with GPTQ residual rounding, via [diffuse-compressor](https://github.com/rootonchair/diffuse-compressor)
83
+ (commit `0965874`). A low-rank BF16 branch absorbs the outliers of each weight and SmoothQuant-style
84
+ scaling migrates activation outliers; GPTQ then rounds the 4-bit residual using calibration statistics.
85
+ Activation scales are dynamic, so nothing is fixed to the calibration inputs.
86
+ - **Calibration data (128 samples per component):** 64 prompts from the qdiff prompt set and 64 image edits
87
+ from the train split of [VyoJ/NHR-Edit-Change_Only](https://huggingface.co/datasets/VyoJ/NHR-Edit-Change_Only).
88
+ Three of every four samples at 512x512, the rest at 1024x1024. The DiT saw 20 denoising steps per sample,
89
+ with the prefix KV cache disabled so the prompt and reference-image tokens pass through every step.
90
+ - **Sensitive layers:** the first and last two DiT blocks and the modulation stay BF16. For the text encoder,
91
+ quantizing attention hurt edits most. With every encoder linear quantized, the "turn the parrots blue"
92
+ recolor below worked in 1 of 8 seeds (rank 32) or 4-5 of 8 (rank 128); quantizing only the MLPs with rank
93
+ 128 brought it to 7 of 8, the same as BF16.
94
+ - **Runtime:** Diffusers' built-in Nunchaku Lite quantizer for the DiT and the same `SVDQW4A4Linear` layers
95
+ for the text encoder, running the [nunchaku-lite](https://github.com/rootonchair/nunchaku-lite) CUDA kernels.
96
+
97
+ The scripts that produced this repository are in `tools/`: `quantize_dit.py`,
98
+ `quantize_text_encoder.py --rank 128 --edge-layers 4 --skip-attention`, `package.py`, `evaluate.py` and
99
+ `parrot_test.py`. They need a checkout of diffuse-compressor (its `examples/` package) at
100
+ `$DIFFUSE_COMPRESSOR`.
101
+
102
+ ## Results
103
+
104
+ RTX PRO 6000 Blackwell (96 GB), Diffusers, 40 steps, guidance 1, no `torch.compile`, after a warmup.
105
+ Peak VRAM is PyTorch's peak allocation.
106
+
107
+ | | BF16 (official) | This repo |
108
+ |---|---|---|
109
+ | Text-to-image, 512x512 | 3.32 s | **1.91 s** (1.74x) |
110
+ | Edit, 512x512 | 3.73 s | **2.18 s** (1.71x) |
111
+ | Text-to-image, 1024x1024 | 14.06 s | **8.12 s** (1.73x) |
112
+ | Peak VRAM, 512x512 | 31.9-32.5 GiB | **18.6-19.3 GiB** |
113
+ | Throughput, 512x512 batch 1-8 | 0.28-0.30 img/s | **0.49-0.55 img/s** |
114
+
115
+ Batching several prompts gives no extra throughput in Diffusers on this GPU; it is saturated at batch 1.
116
+ The pipeline shares one reference-image list across a batch, so edits with different images run one at a time.
117
+
118
+ Closeness to BF16 (LPIPS with AlexNet, same seeds, 512x512; lower is closer, around 0.1 is hard to tell apart):
119
+
120
+ | Set | LPIPS vs BF16 |
121
+ |---|---|
122
+ | 8 text-to-image prompts (portraits, typography, counting, scenes) | 0.148 |
123
+ | 12 held-out NHR-Edit test edits + 1 parrot recolor | 0.046 |
124
+
125
+ Instruction following on a hard recolor ("Turn the parrots blue", 8 seeds, tiles with at least one blue parrot),
126
+ with this repo's DiT:
127
+
128
+ | Text encoder | Followed |
129
+ |---|---|
130
+ | BF16 | 7 / 8 |
131
+ | NVFP4, every linear, rank 32 | 1 / 8 |
132
+ | NVFP4, every linear, rank 128, 4 BF16 edge layers | 4-5 / 8 |
133
+ | **NVFP4, MLPs only, rank 128, 4 BF16 edge layers (this repo)** | **7 / 8** |
134
+
135
+ With the BF16 text encoder, this repo's DiT followed the edit in 8 of 8 seeds, the same as the BF16 DiT.
136
 
137
  ## Limitations
138
 
139
+ - Measured on one GPU model and one software stack; quantization shifts details of individual images.
140
+ - The evaluation is small (8 prompts, 13 edits, one 8-seed recolor test). It is not a benchmark.
141
+ - The kernel build covers `sm_120a`, PyTorch 2.13, CUDA 13 and CPython 3.12 only.
142
  - Non-commercial research and evaluation use only.
143
+
144
+ ## Previous version
145
+
146
+ Until 2026-09-25 this repository held a different combination for SGLang Diffusion: the ModelOpt NVFP4 DiT from
147
+ HangGlidersRule/Darkstar-Qwen-Image-2.1-Base-ModelOpt-W4A4-NVFP4 and the NVFP4 Qwen3-VL encoder from
148
+ BennyDaBall/Qwen-Image-2.1-NVFP4. It is still available in this repository's commit history. It was
149
+ replaced because its 4-bit text encoder lost hard edits and it did not load in Diffusers.
150
 
151
  ## Credits and license
152
 
153
+ - Qwen team: Qwen-Image-2.1.
154
+ - [SVDQuant](https://arxiv.org/abs/2411.05007) (MIT Han Lab), [nunchaku-lite](https://github.com/rootonchair/nunchaku-lite)
155
+ and [diffuse-compressor](https://github.com/rootonchair/diffuse-compressor) (rootonchair).
156
+ - Calibration edits: [VyoJ/NHR-Edit-Change_Only](https://huggingface.co/datasets/VyoJ/NHR-Edit-Change_Only).
157
 
158
+ Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory
159
+ Technology Co., Ltd. All Rights Reserved. See `LICENSE` and `NOTICE`, which lists every modified file.
model_index.json CHANGED
@@ -10,15 +10,15 @@
10
  "FlowMatchEulerDiscreteScheduler"
11
  ],
12
  "text_encoder": [
13
- "transformers",
14
- "Qwen3VLForConditionalGeneration"
15
  ],
16
  "transformer": [
17
- "diffusers",
18
- "QwenImage21Transformer2DModel"
19
  ],
20
  "vae": [
21
  "diffusers",
22
  "AutoencoderKLQwenImage21"
23
  ]
24
- }
 
10
  "FlowMatchEulerDiscreteScheduler"
11
  ],
12
  "text_encoder": [
13
+ "modeling_nunchaku_qwen3vl",
14
+ "NunchakuQwen3VLForConditionalGeneration"
15
  ],
16
  "transformer": [
17
+ "modeling_nunchaku_qwenimage21",
18
+ "NunchakuQwenImage21Transformer2DModel"
19
  ],
20
  "vae": [
21
  "diffusers",
22
  "AutoencoderKLQwenImage21"
23
  ]
24
+ }
text_encoder/config.json CHANGED
@@ -5,6 +5,101 @@
5
  "dtype": "bfloat16",
6
  "image_token_id": 151655,
7
  "model_type": "qwen3_vl",
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8
  "text_config": {
9
  "attention_bias": false,
10
  "attention_dropout": 0.0,
 
5
  "dtype": "bfloat16",
6
  "image_token_id": 151655,
7
  "model_type": "qwen3_vl",
8
+ "nunchaku_lite": {
9
+ "compute_dtype": "bfloat16",
10
+ "quant_method": "nunchaku_lite",
11
+ "svdq_w4a4": {
12
+ "group_size": 16,
13
+ "precision": "nvfp4",
14
+ "rank": 128,
15
+ "targets": [
16
+ "model.language_model.layers.4.mlp.gate_proj",
17
+ "model.language_model.layers.4.mlp.up_proj",
18
+ "model.language_model.layers.4.mlp.down_proj",
19
+ "model.language_model.layers.5.mlp.gate_proj",
20
+ "model.language_model.layers.5.mlp.up_proj",
21
+ "model.language_model.layers.5.mlp.down_proj",
22
+ "model.language_model.layers.6.mlp.gate_proj",
23
+ "model.language_model.layers.6.mlp.up_proj",
24
+ "model.language_model.layers.6.mlp.down_proj",
25
+ "model.language_model.layers.7.mlp.gate_proj",
26
+ "model.language_model.layers.7.mlp.up_proj",
27
+ "model.language_model.layers.7.mlp.down_proj",
28
+ "model.language_model.layers.8.mlp.gate_proj",
29
+ "model.language_model.layers.8.mlp.up_proj",
30
+ "model.language_model.layers.8.mlp.down_proj",
31
+ "model.language_model.layers.9.mlp.gate_proj",
32
+ "model.language_model.layers.9.mlp.up_proj",
33
+ "model.language_model.layers.9.mlp.down_proj",
34
+ "model.language_model.layers.10.mlp.gate_proj",
35
+ "model.language_model.layers.10.mlp.up_proj",
36
+ "model.language_model.layers.10.mlp.down_proj",
37
+ "model.language_model.layers.11.mlp.gate_proj",
38
+ "model.language_model.layers.11.mlp.up_proj",
39
+ "model.language_model.layers.11.mlp.down_proj",
40
+ "model.language_model.layers.12.mlp.gate_proj",
41
+ "model.language_model.layers.12.mlp.up_proj",
42
+ "model.language_model.layers.12.mlp.down_proj",
43
+ "model.language_model.layers.13.mlp.gate_proj",
44
+ "model.language_model.layers.13.mlp.up_proj",
45
+ "model.language_model.layers.13.mlp.down_proj",
46
+ "model.language_model.layers.14.mlp.gate_proj",
47
+ "model.language_model.layers.14.mlp.up_proj",
48
+ "model.language_model.layers.14.mlp.down_proj",
49
+ "model.language_model.layers.15.mlp.gate_proj",
50
+ "model.language_model.layers.15.mlp.up_proj",
51
+ "model.language_model.layers.15.mlp.down_proj",
52
+ "model.language_model.layers.16.mlp.gate_proj",
53
+ "model.language_model.layers.16.mlp.up_proj",
54
+ "model.language_model.layers.16.mlp.down_proj",
55
+ "model.language_model.layers.17.mlp.gate_proj",
56
+ "model.language_model.layers.17.mlp.up_proj",
57
+ "model.language_model.layers.17.mlp.down_proj",
58
+ "model.language_model.layers.18.mlp.gate_proj",
59
+ "model.language_model.layers.18.mlp.up_proj",
60
+ "model.language_model.layers.18.mlp.down_proj",
61
+ "model.language_model.layers.19.mlp.gate_proj",
62
+ "model.language_model.layers.19.mlp.up_proj",
63
+ "model.language_model.layers.19.mlp.down_proj",
64
+ "model.language_model.layers.20.mlp.gate_proj",
65
+ "model.language_model.layers.20.mlp.up_proj",
66
+ "model.language_model.layers.20.mlp.down_proj",
67
+ "model.language_model.layers.21.mlp.gate_proj",
68
+ "model.language_model.layers.21.mlp.up_proj",
69
+ "model.language_model.layers.21.mlp.down_proj",
70
+ "model.language_model.layers.22.mlp.gate_proj",
71
+ "model.language_model.layers.22.mlp.up_proj",
72
+ "model.language_model.layers.22.mlp.down_proj",
73
+ "model.language_model.layers.23.mlp.gate_proj",
74
+ "model.language_model.layers.23.mlp.up_proj",
75
+ "model.language_model.layers.23.mlp.down_proj",
76
+ "model.language_model.layers.24.mlp.gate_proj",
77
+ "model.language_model.layers.24.mlp.up_proj",
78
+ "model.language_model.layers.24.mlp.down_proj",
79
+ "model.language_model.layers.25.mlp.gate_proj",
80
+ "model.language_model.layers.25.mlp.up_proj",
81
+ "model.language_model.layers.25.mlp.down_proj",
82
+ "model.language_model.layers.26.mlp.gate_proj",
83
+ "model.language_model.layers.26.mlp.up_proj",
84
+ "model.language_model.layers.26.mlp.down_proj",
85
+ "model.language_model.layers.27.mlp.gate_proj",
86
+ "model.language_model.layers.27.mlp.up_proj",
87
+ "model.language_model.layers.27.mlp.down_proj",
88
+ "model.language_model.layers.28.mlp.gate_proj",
89
+ "model.language_model.layers.28.mlp.up_proj",
90
+ "model.language_model.layers.28.mlp.down_proj",
91
+ "model.language_model.layers.29.mlp.gate_proj",
92
+ "model.language_model.layers.29.mlp.up_proj",
93
+ "model.language_model.layers.29.mlp.down_proj",
94
+ "model.language_model.layers.30.mlp.gate_proj",
95
+ "model.language_model.layers.30.mlp.up_proj",
96
+ "model.language_model.layers.30.mlp.down_proj",
97
+ "model.language_model.layers.31.mlp.gate_proj",
98
+ "model.language_model.layers.31.mlp.up_proj",
99
+ "model.language_model.layers.31.mlp.down_proj"
100
+ ]
101
+ }
102
+ },
103
  "text_config": {
104
  "attention_bias": false,
105
  "attention_dropout": 0.0,
text_encoder/model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:42882503fe37fbc72f374fb9e55e79f020f9cecca02463417176fdd902be470e
3
- size 7549900080
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:83c4462b5ce6b98e2708ca53031e2f9a60a8e402c4c4a0ccd19cb2df8e817ae2
3
+ size 11812006400
text_encoder/modeling_nunchaku_qwen3vl.py ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Qwen3-VL text encoder with SVDQuant NVFP4 (Nunchaku Lite) linears.
2
+
3
+ Loaded by Diffusers as a custom pipeline component (`trust_remote_code=True`).
4
+ `config.json` holds the compact Nunchaku Lite config under `nunchaku_lite`; the
5
+ listed linears are swapped for Diffusers' own `SVDQW4A4Linear` before the
6
+ quantized state dict is loaded, so inference uses the same NVFP4 kernels as the
7
+ transformer. Everything else (vision tower, embeddings, the BF16 layers) loads
8
+ as in the stock model.
9
+ """
10
+
11
+ from pathlib import Path
12
+
13
+ import torch
14
+ from accelerate import init_empty_weights
15
+ from huggingface_hub import snapshot_download
16
+
17
+ from .nunchaku_kernels import KERNELS # noqa: F401 (sets up the kernels before Diffusers imports them)
18
+
19
+ # isort: off -- must stay below `.nunchaku_kernels`, which prepares the kernels this import loads
20
+ from diffusers.quantizers.nunchaku.utils import replace_with_nunchaku_linear # noqa: E402
21
+ from safetensors.torch import load_file # noqa: E402
22
+ from transformers import Qwen3VLForConditionalGeneration # noqa: E402
23
+ # isort: on
24
+
25
+
26
+ class NunchakuQwen3VLForConditionalGeneration(Qwen3VLForConditionalGeneration):
27
+ @classmethod
28
+ def from_pretrained(cls, pretrained_model_name_or_path, *args, subfolder: str = "", **kwargs):
29
+ path = Path(pretrained_model_name_or_path) / subfolder
30
+ if not path.is_dir():
31
+ path = Path(snapshot_download(str(pretrained_model_name_or_path), allow_patterns=[f"{subfolder}/*"])) / subfolder
32
+ dtype = kwargs.get("dtype") or kwargs.get("torch_dtype") or torch.bfloat16
33
+ if dtype == "auto":
34
+ dtype = torch.bfloat16
35
+
36
+ config = cls.config_class.from_pretrained(path)
37
+ with init_empty_weights(): # parameters on meta, buffers (e.g. rotary tables) materialized
38
+ model = cls(config)
39
+ replace_with_nunchaku_linear(model, config.nunchaku_lite, dtype)
40
+ # strict=False only because non-persistent buffers are absent; both checks below stay strict
41
+ result = model.load_state_dict(load_file(path / "model.safetensors"), strict=False, assign=True)
42
+ if result.unexpected_keys:
43
+ raise ValueError(f"Checkpoint has {len(result.unexpected_keys)} unexpected tensors, e.g. {result.unexpected_keys[:3]}")
44
+ missing = [name for name, p in model.named_parameters() if p.device.type == "meta"]
45
+ if missing:
46
+ raise ValueError(f"Checkpoint is missing {len(missing)} parameters, e.g. {missing[:3]}")
47
+ return model.eval()
text_encoder/nunchaku_kernels.py ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Shared by text_encoder/ and transformer/: whichever component Diffusers loads first sets this up.
2
+ #
3
+ # Diffusers fetches the Nunchaku Lite kernels from `rootonchair/nunchaku-lite-kernels`, which is
4
+ # no longer downloadable. Loading this repo with trust_remote_code=True already runs this code,
5
+ # so point that kernel name at the rebuilt copy before Diffusers imports its Nunchaku utilities.
6
+ import os
7
+
8
+ from huggingface_hub import snapshot_download
9
+
10
+ KERNELS = "rootonchair/nunchaku-lite-kernels"
11
+ if KERNELS not in os.environ.get("LOCAL_KERNELS", ""):
12
+ local = f"{KERNELS}={snapshot_download('joseplcam/nunchaku-lite-kernels')}"
13
+ os.environ["LOCAL_KERNELS"] = ":".join(filter(None, [os.environ.get("LOCAL_KERNELS"), local]))
14
+ if "DIFFUSERS_TRUST_REMOTE_KERNELS" not in os.environ:
15
+ # Diffusers read this variable into a constant at import time, so set both.
16
+ import diffusers.utils.constants
17
+
18
+ os.environ["DIFFUSERS_TRUST_REMOTE_KERNELS"] = "true"
19
+ diffusers.utils.constants.DIFFUSERS_TRUST_REMOTE_KERNELS = True
tools/evaluate.py ADDED
@@ -0,0 +1,103 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Compare a Qwen-Image-2.1 pipeline against the BF16 original on text-to-image and edits.
2
+
3
+ Writes, per pipeline: images, per-item LPIPS vs BF16 (same seeds), timings and peak VRAM.
4
+ Edits come from the NHR-Edit test split (held out from calibration) plus the parrot
5
+ recolor that 4-bit models failed before. Instruction following is judged by eye from
6
+ the saved grids; LPIPS only measures closeness to BF16.
7
+
8
+ Usage:
9
+ python tools/qwen21_nvfp4/evaluate.py <pipeline path or id> <tag> [--trust-remote-code] [--baseline <bf16 out dir>]
10
+ """
11
+
12
+ import argparse
13
+ import json
14
+ import time
15
+ import urllib.request
16
+ from io import BytesIO
17
+ from pathlib import Path
18
+
19
+ import datasets
20
+ import torch
21
+ from diffusers import DiffusionPipeline
22
+ from PIL import Image
23
+
24
+ PROMPTS = [
25
+ "A capybara reading a book by candlelight",
26
+ "A red vintage bicycle leaning against a blue door",
27
+ "A bowl of ramen with a soft boiled egg, top view",
28
+ "A neon sign that says OPEN in a rainy alley",
29
+ "A glass teapot with blooming flower tea on a table",
30
+ "A portrait of an elderly fisherman with a knitted cap, soft window light",
31
+ 'A chalkboard menu that reads "SOUP OF THE DAY: TOMATO"',
32
+ "Three red apples and two green pears on a wooden table",
33
+ ]
34
+ PARROTS = "https://r0k.us/graphics/kodak/kodak/kodim23.png"
35
+
36
+
37
+ def edit_items(n: int) -> list[tuple[str, Image.Image, str]]:
38
+ data = datasets.load_dataset("VyoJ/NHR-Edit-Change_Only", split="test")
39
+ items = [(f"nhr{i}", data[i]["source"].convert("RGB"), data[i]["edit_instruction"]) for i in range(n)]
40
+ parrots = Image.open(BytesIO(urllib.request.urlopen(PARROTS).read())).convert("RGB")
41
+ return items + [("parrots", parrots, "Turn the parrots blue")]
42
+
43
+
44
+ def main():
45
+ parser = argparse.ArgumentParser()
46
+ parser.add_argument("pipeline")
47
+ parser.add_argument("tag")
48
+ parser.add_argument("--trust-remote-code", action="store_true")
49
+ parser.add_argument("--edits", type=int, default=12)
50
+ parser.add_argument("--size", type=int, default=512)
51
+ parser.add_argument("--out", default="outputs/qwen21_eval")
52
+ parser.add_argument("--baseline", help="output dir of the BF16 run, for LPIPS")
53
+ args = parser.parse_args()
54
+
55
+ out = Path(args.out) / args.tag
56
+ out.mkdir(parents=True, exist_ok=True)
57
+ pipe = DiffusionPipeline.from_pretrained(args.pipeline, dtype=torch.bfloat16, trust_remote_code=args.trust_remote_code).to("cuda")
58
+ pipe.set_progress_bar_config(disable=True)
59
+ gen = lambda seed: torch.Generator("cuda").manual_seed(seed) # noqa: E731
60
+
61
+ def run(name, seed, **kwargs):
62
+ torch.cuda.synchronize()
63
+ start = time.perf_counter()
64
+ image = pipe(num_inference_steps=40, generator=gen(seed), **kwargs).images[0]
65
+ torch.cuda.synchronize()
66
+ image.save(out / f"{name}.png")
67
+ return time.perf_counter() - start
68
+
69
+ run("warmup", 0, prompt="warmup", height=args.size, width=args.size)
70
+ run("warmup", 1, prompt="warmup", image=Image.open(out / "warmup.png"), output_resolution=args.size)
71
+ torch.cuda.reset_peak_memory_stats()
72
+ times = {"t2i": [], "edit": []}
73
+ for i, prompt in enumerate(PROMPTS):
74
+ times["t2i"].append(run(f"t2i{i}", 1000 + i, prompt=prompt, height=args.size, width=args.size))
75
+ for name, image, prompt in edit_items(args.edits):
76
+ times["edit"].append(run(name, 2000, prompt=prompt, image=image, output_resolution=args.size))
77
+ report = {
78
+ "tag": args.tag,
79
+ "size": args.size,
80
+ "t2i_s": round(sum(times["t2i"]) / len(times["t2i"]), 3),
81
+ "edit_s": round(sum(times["edit"]) / len(times["edit"]), 3),
82
+ "peak_vram_gib": round(torch.cuda.max_memory_allocated() / 2**30, 2),
83
+ }
84
+ if args.baseline:
85
+ import lpips
86
+ import numpy as np
87
+
88
+ metric = lpips.LPIPS(net="alex", verbose=False)
89
+ load = lambda p: torch.from_numpy(np.array(Image.open(p).convert("RGB"))).permute(2, 0, 1)[None].float() / 127.5 - 1 # noqa: E731
90
+ scores = {
91
+ p.stem: round(metric(load(Path(args.baseline) / p.name), load(p)).item(), 4)
92
+ for p in sorted(out.glob("*.png"))
93
+ if p.stem != "warmup" and (Path(args.baseline) / p.name).exists()
94
+ }
95
+ report["lpips"] = scores
96
+ report["lpips_t2i_mean"] = round(sum(v for k, v in scores.items() if k.startswith("t2i")) / len(PROMPTS), 4)
97
+ report["lpips_edit_mean"] = round(sum(v for k, v in scores.items() if not k.startswith("t2i")) / (args.edits + 1), 4)
98
+ (out / "report.json").write_text(json.dumps(report, indent=1))
99
+ print(json.dumps(report))
100
+
101
+
102
+ if __name__ == "__main__":
103
+ main()
tools/modeling_nunchaku_qwen3vl.py ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Qwen3-VL text encoder with SVDQuant NVFP4 (Nunchaku Lite) linears.
2
+
3
+ Loaded by Diffusers as a custom pipeline component (`trust_remote_code=True`).
4
+ `config.json` holds the compact Nunchaku Lite config under `nunchaku_lite`; the
5
+ listed linears are swapped for Diffusers' own `SVDQW4A4Linear` before the
6
+ quantized state dict is loaded, so inference uses the same NVFP4 kernels as the
7
+ transformer. Everything else (vision tower, embeddings, the BF16 layers) loads
8
+ as in the stock model.
9
+ """
10
+
11
+ from pathlib import Path
12
+
13
+ import torch
14
+ from accelerate import init_empty_weights
15
+ from huggingface_hub import snapshot_download
16
+
17
+ from .nunchaku_kernels import KERNELS # noqa: F401 (sets up the kernels before Diffusers imports them)
18
+
19
+ # isort: off -- must stay below `.nunchaku_kernels`, which prepares the kernels this import loads
20
+ from diffusers.quantizers.nunchaku.utils import replace_with_nunchaku_linear # noqa: E402
21
+ from safetensors.torch import load_file # noqa: E402
22
+ from transformers import Qwen3VLForConditionalGeneration # noqa: E402
23
+ # isort: on
24
+
25
+
26
+ class NunchakuQwen3VLForConditionalGeneration(Qwen3VLForConditionalGeneration):
27
+ @classmethod
28
+ def from_pretrained(cls, pretrained_model_name_or_path, *args, subfolder: str = "", **kwargs):
29
+ path = Path(pretrained_model_name_or_path) / subfolder
30
+ if not path.is_dir():
31
+ path = Path(snapshot_download(str(pretrained_model_name_or_path), allow_patterns=[f"{subfolder}/*"])) / subfolder
32
+ dtype = kwargs.get("dtype") or kwargs.get("torch_dtype") or torch.bfloat16
33
+ if dtype == "auto":
34
+ dtype = torch.bfloat16
35
+
36
+ config = cls.config_class.from_pretrained(path)
37
+ with init_empty_weights(): # parameters on meta, buffers (e.g. rotary tables) materialized
38
+ model = cls(config)
39
+ replace_with_nunchaku_linear(model, config.nunchaku_lite, dtype)
40
+ # strict=False only because non-persistent buffers are absent; both checks below stay strict
41
+ result = model.load_state_dict(load_file(path / "model.safetensors"), strict=False, assign=True)
42
+ if result.unexpected_keys:
43
+ raise ValueError(f"Checkpoint has {len(result.unexpected_keys)} unexpected tensors, e.g. {result.unexpected_keys[:3]}")
44
+ missing = [name for name, p in model.named_parameters() if p.device.type == "meta"]
45
+ if missing:
46
+ raise ValueError(f"Checkpoint is missing {len(missing)} parameters, e.g. {missing[:3]}")
47
+ return model.eval()
tools/modeling_nunchaku_qwenimage21.py ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Qwen-Image-2.1 transformer, unchanged, loaded as a custom component.
2
+
3
+ Its only job is importing `nunchaku_kernels`, so the NVFP4 kernels are set up even when
4
+ Diffusers loads this component before the text encoder. The Nunchaku Lite quantization
5
+ itself is handled by Diffusers' built-in quantizer from `config.json`.
6
+ """
7
+
8
+ from diffusers import QwenImage21Transformer2DModel
9
+
10
+ from .nunchaku_kernels import KERNELS # noqa: F401
11
+
12
+
13
+ class NunchakuQwenImage21Transformer2DModel(QwenImage21Transformer2DModel):
14
+ pass
tools/nunchaku_kernels.py ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Shared by text_encoder/ and transformer/: whichever component Diffusers loads first sets this up.
2
+ #
3
+ # Diffusers fetches the Nunchaku Lite kernels from `rootonchair/nunchaku-lite-kernels`, which is
4
+ # no longer downloadable. Loading this repo with trust_remote_code=True already runs this code,
5
+ # so point that kernel name at the rebuilt copy before Diffusers imports its Nunchaku utilities.
6
+ import os
7
+
8
+ from huggingface_hub import snapshot_download
9
+
10
+ KERNELS = "rootonchair/nunchaku-lite-kernels"
11
+ if KERNELS not in os.environ.get("LOCAL_KERNELS", ""):
12
+ local = f"{KERNELS}={snapshot_download('joseplcam/nunchaku-lite-kernels')}"
13
+ os.environ["LOCAL_KERNELS"] = ":".join(filter(None, [os.environ.get("LOCAL_KERNELS"), local]))
14
+ if "DIFFUSERS_TRUST_REMOTE_KERNELS" not in os.environ:
15
+ # Diffusers read this variable into a constant at import time, so set both.
16
+ import diffusers.utils.constants
17
+
18
+ os.environ["DIFFUSERS_TRUST_REMOTE_KERNELS"] = "true"
19
+ diffusers.utils.constants.DIFFUSERS_TRUST_REMOTE_KERNELS = True
tools/package.py ADDED
@@ -0,0 +1,81 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Package the quantized DiT and text encoder into one Diffusers pipeline folder.
2
+
3
+ - transformer/: Nunchaku Lite NVFP4 checkpoint (loaded by Diffusers' built-in quantizer)
4
+ - text_encoder/: Nunchaku Lite NVFP4 checkpoint + custom component class (trust_remote_code)
5
+ - vae/: official weights stored in BF16
6
+ - processor/, scheduler/, model_index.json, LICENSE: from Qwen/Qwen-Image-2.1
7
+
8
+ Usage:
9
+ python tools/qwen21_nvfp4/package.py --dit <dit.safetensors> --text-encoder <te.safetensors> --out <dir>
10
+ """
11
+
12
+ import argparse
13
+ import json
14
+ import os
15
+ import shutil
16
+ import sys
17
+ from pathlib import Path
18
+
19
+ import torch
20
+ from diffusers import AutoencoderKLQwenImage21
21
+
22
+ # examples/ is not part of the installed package, so a checkout of the repo is needed
23
+ sys.path.insert(0, os.environ.get("DIFFUSE_COMPRESSOR", str(Path.home() / "src" / "diffuse-compressor")))
24
+ from examples.convert_nunchaku_lite_diffusers import ( # noqa: E402
25
+ build_diffusers_quantization_config,
26
+ package_diffusers_pipeline,
27
+ )
28
+
29
+ HERE = Path(__file__).parent
30
+ # component -> (module file, class); both import nunchaku_kernels, which ships next to each
31
+ CUSTOM_COMPONENTS = {
32
+ "text_encoder": ("modeling_nunchaku_qwen3vl", "NunchakuQwen3VLForConditionalGeneration"),
33
+ "transformer": ("modeling_nunchaku_qwenimage21", "NunchakuQwenImage21Transformer2DModel"),
34
+ }
35
+
36
+
37
+ def main():
38
+ parser = argparse.ArgumentParser()
39
+ parser.add_argument("--dit", required=True)
40
+ parser.add_argument("--text-encoder", required=True)
41
+ parser.add_argument("--out", required=True)
42
+ parser.add_argument("--model-id", default="Qwen/Qwen-Image-2.1")
43
+ args = parser.parse_args()
44
+
45
+ out = package_diffusers_pipeline(args.dit, args.model_id, args.out, compute_dtype="bfloat16")
46
+
47
+ te = out / "text_encoder"
48
+ for dense in [*te.glob("*.safetensors"), *te.glob("*.safetensors.index.json")]:
49
+ dense.unlink()
50
+ shutil.copy2(args.text_encoder, te / "model.safetensors")
51
+ config = json.loads((te / "config.json").read_text())
52
+ config["nunchaku_lite"] = build_diffusers_quantization_config(args.text_encoder, compute_dtype="bfloat16")
53
+ (te / "config.json").write_text(json.dumps(config, indent=2, sort_keys=True) + "\n")
54
+
55
+ index = json.loads((out / "model_index.json").read_text())
56
+ for component, (module, cls) in CUSTOM_COMPONENTS.items():
57
+ shutil.copy2(HERE / f"{module}.py", out / component / f"{module}.py")
58
+ shutil.copy2(HERE / "nunchaku_kernels.py", out / component / "nunchaku_kernels.py")
59
+ index[component] = [module, cls]
60
+ (out / "model_index.json").write_text(json.dumps(index, indent=2) + "\n")
61
+
62
+ vae = AutoencoderKLQwenImage21.from_pretrained(out / "vae", dtype=torch.bfloat16)
63
+ shutil.rmtree(out / "vae")
64
+ vae.save_pretrained(out / "vae")
65
+ vae_config = json.loads((out / "vae" / "config.json").read_text())
66
+ vae_config.pop("_name_or_path", None) # a local path, meaningless on the Hub
67
+ (out / "vae" / "config.json").write_text(json.dumps(vae_config, indent=2, sort_keys=True) + "\n")
68
+
69
+ for leftover in ("assets", ".gitattributes"):
70
+ path = out / leftover
71
+ shutil.rmtree(path) if path.is_dir() else path.unlink(missing_ok=True)
72
+ for doc in ("README.md", "NOTICE"): # model card and license notice
73
+ shutil.copy2(HERE / doc, out / doc)
74
+ (out / "tools").mkdir()
75
+ for script in sorted(HERE.glob("*.py")): # the recipe, for reproducibility
76
+ shutil.copy2(script, out / "tools" / script.name)
77
+ print(out)
78
+
79
+
80
+ if __name__ == "__main__":
81
+ main()
tools/parrot_test.py ADDED
@@ -0,0 +1,69 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Hard-edit check: "Turn the parrots blue" over 8 seeds at 512x512, per text encoder variant.
2
+
3
+ 4-bit text encoders tend to lose this recolor (the model keeps the reference colors), so it
4
+ separates encoder variants that LPIPS alone does not. Each row of the output grid is one
5
+ variant; count the tiles with at least one blue parrot.
6
+
7
+ Usage:
8
+ python tools/qwen21_nvfp4/parrot_test.py <packaged pipeline> <out.jpg> bf16 <encoder checkpoint> ...
9
+ """
10
+
11
+ import io
12
+ import json
13
+ import shutil
14
+ import sys
15
+ import tempfile
16
+ import urllib.request
17
+ from pathlib import Path
18
+
19
+ import torch
20
+ from diffusers import DiffusionPipeline
21
+ from PIL import Image
22
+ from transformers import Qwen3VLForConditionalGeneration
23
+
24
+ sys.path.insert(0, str(Path(__file__).parent))
25
+ from package import build_diffusers_quantization_config # noqa: E402
26
+
27
+ PARROTS = "https://r0k.us/graphics/kodak/kodak/kodim23.png"
28
+
29
+
30
+ def load_encoder(encoder_cls, pipeline_dir: Path, variant: str):
31
+ if variant == "bf16":
32
+ return Qwen3VLForConditionalGeneration.from_pretrained("Qwen/Qwen-Image-2.1", subfolder="text_encoder", dtype=torch.bfloat16)
33
+ with tempfile.TemporaryDirectory() as tmp: # the packaged text_encoder folder with this checkpoint swapped in
34
+ folder = Path(tmp)
35
+ for name in ("generation_config.json",):
36
+ shutil.copy2(pipeline_dir / "text_encoder" / name, folder / name)
37
+ config = json.loads((pipeline_dir / "text_encoder" / "config.json").read_text())
38
+ config["nunchaku_lite"] = build_diffusers_quantization_config(variant, compute_dtype="bfloat16")
39
+ (folder / "config.json").write_text(json.dumps(config))
40
+ (folder / "model.safetensors").symlink_to(Path(variant).resolve())
41
+ return encoder_cls.from_pretrained(folder, dtype=torch.bfloat16)
42
+
43
+
44
+ def main():
45
+ pipeline_dir, out, variants = Path(sys.argv[1]), sys.argv[2], sys.argv[3:]
46
+ pipe = DiffusionPipeline.from_pretrained(pipeline_dir, dtype=torch.bfloat16, trust_remote_code=True).to("cuda")
47
+ pipe.set_progress_bar_config(disable=True)
48
+ parrots = Image.open(io.BytesIO(urllib.request.urlopen(PARROTS).read())).convert("RGB")
49
+ encoder_cls = type(pipe.text_encoder)
50
+ rows = []
51
+ for variant in variants:
52
+ pipe.text_encoder = None
53
+ torch.cuda.empty_cache()
54
+ pipe.text_encoder = load_encoder(encoder_cls, pipeline_dir, variant).to("cuda")
55
+ rows.append([
56
+ pipe(prompt="Turn the parrots blue", image=parrots, output_resolution=512, num_inference_steps=40,
57
+ generator=torch.Generator("cuda").manual_seed(seed)).images[0].convert("RGB").resize((128, 128))
58
+ for seed in range(1, 9)
59
+ ])
60
+ print(Path(variant).name, "done", flush=True)
61
+ grid = Image.new("RGB", (128 * 8, 128 * len(rows)))
62
+ for y, row in enumerate(rows):
63
+ for x, tile in enumerate(row):
64
+ grid.paste(tile, (128 * x, 128 * y))
65
+ grid.save(out)
66
+
67
+
68
+ if __name__ == "__main__":
69
+ main()
tools/quantize_dit.py ADDED
@@ -0,0 +1,144 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """SVDQuant + GPTQ NVFP4 quantization of the Qwen-Image-2.1 DiT with diffuse-compressor.
2
+
3
+ Calibration mixes text-to-image prompts and image edits (NHR-Edit train split),
4
+ mostly at 512x512 with some 1024x1024, so both workloads shape the scales.
5
+ The first two and last two blocks and the global modulation stay in BF16.
6
+
7
+ Usage (from the repo root, envs/base_final interpreter):
8
+ python tools/qwen21_nvfp4/quantize_dit.py --num-samples 128 --rank 32
9
+ """
10
+
11
+ import argparse
12
+ import os
13
+ import random
14
+ import sys
15
+ from dataclasses import replace
16
+ from pathlib import Path
17
+
18
+ import torch
19
+ from diffusers import QwenImage21Pipeline
20
+
21
+ # examples/ is not part of the installed package, so a checkout of the repo is needed
22
+ sys.path.insert(0, os.environ.get("DIFFUSE_COMPRESSOR", str(Path.home() / "src" / "diffuse-compressor")))
23
+
24
+ from diffuse_compressor import ( # noqa: E402
25
+ CalibrationSpec,
26
+ ExportSpec,
27
+ LoggingConfig,
28
+ QuantizationCacheSpec,
29
+ quantize_and_export,
30
+ )
31
+ from diffuse_compressor.config import GptqSpec # noqa: E402
32
+ from examples.image_to_image.quantize_hf import scan_linear_targets # noqa: E402
33
+ from examples.text_to_image.utils import ( # noqa: E402
34
+ make_generator,
35
+ save_diffusers_images,
36
+ standard_prompt_records,
37
+ svdquant_spec,
38
+ )
39
+
40
+ SKIP = [
41
+ "transformer_blocks.0.*",
42
+ "transformer_blocks.1.*",
43
+ "transformer_blocks.30.*",
44
+ "transformer_blocks.31.*",
45
+ "modulation*",
46
+ ]
47
+
48
+
49
+ def records(num_samples: int, edit_dataset: str) -> list[dict]:
50
+ import datasets
51
+
52
+ half = num_samples // 2
53
+ rng = random.Random(0)
54
+ # 3 of every 4 samples at 512x512 (the working resolution), the rest at 1024x1024
55
+ size = lambda i: 1024 if i % 4 == 3 else 512 # noqa: E731
56
+
57
+ t2i = [
58
+ {"filename": r["filename"], "prompt": r["prompt"], "seed": r["seed"], "size": size(i)}
59
+ for i, r in enumerate(standard_prompt_records(num_samples - half))
60
+ ]
61
+ data = datasets.load_dataset(edit_dataset, split="train")
62
+ edits = [
63
+ {
64
+ "filename": f"edit-{j}",
65
+ "prompt": data[j]["edit_instruction"],
66
+ "image": data[j]["source"].convert("RGB"),
67
+ "seed": 100_000 + j,
68
+ "size": size(i),
69
+ }
70
+ for i, j in enumerate(sorted(rng.sample(range(len(data)), half)))
71
+ ]
72
+ return t2i + edits
73
+
74
+
75
+ def forward_fn(pipe, steps: int, device: str):
76
+ def forward(sample: dict):
77
+ kwargs = {
78
+ "prompt": sample["prompt"],
79
+ "num_inference_steps": steps,
80
+ "generator": make_generator(sample["seed"], device=device),
81
+ # the prefix KV cache is a stateful object that calibration replay cannot move or
82
+ # re-run; without it the prefix tokens pass through the same linears every step
83
+ "use_kv_cache": False,
84
+ }
85
+ if sample.get("image") is not None:
86
+ kwargs.update(image=sample["image"], output_resolution=sample["size"])
87
+ else:
88
+ kwargs.update(height=sample["size"], width=sample["size"])
89
+ return pipe(**kwargs)
90
+
91
+ return forward
92
+
93
+
94
+ def main():
95
+ parser = argparse.ArgumentParser()
96
+ parser.add_argument("--model-id", default="Qwen/Qwen-Image-2.1")
97
+ parser.add_argument("--edit-dataset", default="VyoJ/NHR-Edit-Change_Only")
98
+ parser.add_argument("--num-samples", type=int, default=128)
99
+ parser.add_argument("--steps", type=int, default=20)
100
+ parser.add_argument("--rank", type=int, default=32)
101
+ parser.add_argument("--no-gptq", action="store_true")
102
+ parser.add_argument("--out-dir", default="outputs/qwen21_nvfp4")
103
+ parser.add_argument("--cache-mode", choices=("reuse", "refresh", "disabled"), default="reuse")
104
+ parser.add_argument("--inspect", action="store_true")
105
+ args = parser.parse_args()
106
+
107
+ device = "cuda"
108
+ pipe = QwenImage21Pipeline.from_pretrained(args.model_id, torch_dtype=torch.bfloat16).to(device)
109
+ pipe.set_progress_bar_config(disable=True)
110
+ scan = scan_linear_targets(pipe.transformer, precision="nvfp4", rank=args.rank, skip=SKIP)
111
+ print(scan.format_text().splitlines()[0])
112
+ if args.inspect:
113
+ print(scan.format_text())
114
+ return
115
+
116
+ tag = f"svdq-nvfp4_r{args.rank}{'' if args.no_gptq else '-gptq'}-qwen-image-2.1"
117
+ out = Path(args.out_dir)
118
+ cache = out / "calibration" / tag
119
+ spec = replace(svdquant_spec("nvfp4", compute_device=device), rank=args.rank, gptq=GptqSpec(enabled=not args.no_gptq))
120
+ quantize_and_export(
121
+ pipe.transformer,
122
+ spec,
123
+ scan.target_config,
124
+ CalibrationSpec(
125
+ samples=records(args.num_samples, args.edit_dataset),
126
+ num_samples=args.num_samples,
127
+ cache_num_samples=args.num_samples,
128
+ batch_size=1,
129
+ cache_dir=cache / "inputs",
130
+ cache_mode=args.cache_mode,
131
+ forward_fn=forward_fn(pipe, args.steps, device),
132
+ max_rows_per_target=4096,
133
+ artifact_cache=QuantizationCacheSpec(cache / "artifacts", args.cache_mode),
134
+ output_dir=cache / "inputs" / "samples",
135
+ output_save_fn=save_diffusers_images,
136
+ scope_capture_mode="all_targets",
137
+ ),
138
+ ExportSpec(output=out / "checkpoints" / f"{tag}.safetensors"),
139
+ LoggingConfig(name=tag),
140
+ )
141
+
142
+
143
+ if __name__ == "__main__":
144
+ main()
tools/quantize_text_encoder.py ADDED
@@ -0,0 +1,121 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """SVDQuant + GPTQ NVFP4 quantization of the Qwen-Image-2.1 text encoder (Qwen3-VL 8B).
2
+
3
+ Only the language-model linears are quantized. The vision tower, embeddings,
4
+ lm_head (unused by the pipeline) and the first/last `--edge-layers` decoder
5
+ layers stay BF16: the pipeline feeds the last decoder layer's output straight
6
+ into the DiT. Calibration uses the same text-to-image and edit records as the DiT, run
7
+ through the full pipeline so condition images are resized exactly as in use.
8
+
9
+ Usage (from the repo root, envs/base_final interpreter):
10
+ python tools/qwen21_nvfp4/quantize_text_encoder.py --num-samples 128 --rank 32
11
+ """
12
+
13
+ import argparse
14
+ import sys
15
+ from dataclasses import replace
16
+ from pathlib import Path
17
+
18
+ import torch
19
+ from diffusers import QwenImage21Pipeline
20
+
21
+ sys.path.insert(0, str(Path(__file__).parent))
22
+ from quantize_dit import records # noqa: E402 (also puts diffuse-compressor on sys.path)
23
+
24
+ from diffuse_compressor import ( # noqa: E402
25
+ CalibrationSpec,
26
+ ExportSpec,
27
+ LoggingConfig,
28
+ QuantizationCacheSpec,
29
+ quantize_and_export,
30
+ )
31
+ from diffuse_compressor.calibration import replay # noqa: E402
32
+ from diffuse_compressor.config import GptqSpec # noqa: E402
33
+ from examples.image_to_image.quantize_hf import scan_linear_targets # noqa: E402
34
+ from examples.text_to_image.utils import make_generator, svdquant_spec # noqa: E402
35
+
36
+ # Full-model replay CPU-offloads the model through Accelerate, which puts the 16 GB encoder in
37
+ # host RAM and trips the 90% RAM guard on a 32 GB machine. The GPU has room for it, and
38
+ # returning False makes the library fall back to `model.to(device)`.
39
+ replay._accelerate_cpu_offload_for_full_replay = lambda *args, **kwargs: False
40
+
41
+ NUM_LAYERS = 36
42
+
43
+
44
+ def skip_patterns(edge_layers: int, skip_attention: bool) -> list[str]:
45
+ """Vision tower, lm_head, the first/last `edge_layers` decoder layers and optionally all attention stay BF16."""
46
+ edges = [*range(edge_layers), *range(NUM_LAYERS - edge_layers, NUM_LAYERS)]
47
+ patterns = ["model.visual.*", "lm_head", *(f"model.language_model.layers.{i}.*" for i in edges)]
48
+ return patterns + (["model.language_model.layers.*.self_attn.*"] if skip_attention else [])
49
+
50
+
51
+ def forward_fn(pipe, device: str):
52
+ def forward(sample: dict):
53
+ kwargs = {
54
+ "prompt": sample["prompt"],
55
+ "num_inference_steps": 1, # the encoder runs once per call; one DiT step is enough
56
+ "output_type": "latent",
57
+ "generator": make_generator(sample["seed"], device=device),
58
+ }
59
+ if sample.get("image") is not None:
60
+ kwargs.update(image=sample["image"], output_resolution=sample["size"])
61
+ else:
62
+ kwargs.update(height=sample["size"], width=sample["size"])
63
+ return pipe(**kwargs)
64
+
65
+ return forward
66
+
67
+
68
+ def main():
69
+ parser = argparse.ArgumentParser()
70
+ parser.add_argument("--model-id", default="Qwen/Qwen-Image-2.1")
71
+ parser.add_argument("--edit-dataset", default="VyoJ/NHR-Edit-Change_Only")
72
+ parser.add_argument("--num-samples", type=int, default=128)
73
+ parser.add_argument("--rank", type=int, default=32)
74
+ parser.add_argument("--edge-layers", type=int, default=2, help="first/last decoder layers kept in BF16")
75
+ parser.add_argument("--skip-attention", action="store_true", help="keep all attention projections in BF16")
76
+ parser.add_argument("--no-gptq", action="store_true")
77
+ parser.add_argument("--out-dir", default="outputs/qwen21_nvfp4")
78
+ parser.add_argument("--cache-mode", choices=("reuse", "refresh", "disabled"), default="reuse")
79
+ parser.add_argument("--inspect", action="store_true")
80
+ args = parser.parse_args()
81
+
82
+ device = "cuda"
83
+ pipe = QwenImage21Pipeline.from_pretrained(args.model_id, dtype=torch.bfloat16).to(device)
84
+ pipe.set_progress_bar_config(disable=True)
85
+ encoder = pipe.text_encoder
86
+ scan = scan_linear_targets(
87
+ encoder, precision="nvfp4", rank=args.rank, skip=skip_patterns(args.edge_layers, args.skip_attention)
88
+ )
89
+ print(scan.format_text().splitlines()[0])
90
+ if args.inspect:
91
+ print(scan.format_text())
92
+ return
93
+
94
+ variant = f"r{args.rank}_e{args.edge_layers}{'_mlp' if args.skip_attention else ''}{'' if args.no_gptq else '-gptq'}"
95
+ tag = f"svdq-nvfp4_{variant}-qwen-image-2.1-text-encoder"
96
+ out = Path(args.out_dir)
97
+ cache = out / "calibration" / tag
98
+ spec = replace(svdquant_spec("nvfp4", compute_device=device), rank=args.rank, gptq=GptqSpec(enabled=not args.no_gptq))
99
+ quantize_and_export(
100
+ encoder,
101
+ spec,
102
+ scan.target_config,
103
+ CalibrationSpec(
104
+ samples=records(args.num_samples, args.edit_dataset),
105
+ num_samples=args.num_samples,
106
+ cache_num_samples=args.num_samples,
107
+ batch_size=1,
108
+ cache_dir=cache / "inputs",
109
+ cache_mode=args.cache_mode,
110
+ forward_fn=forward_fn(pipe, device),
111
+ max_rows_per_target=4096,
112
+ artifact_cache=QuantizationCacheSpec(cache / "artifacts", args.cache_mode),
113
+ scope_capture_mode="all_targets",
114
+ ),
115
+ ExportSpec(output=out / "checkpoints" / f"{tag}.safetensors"),
116
+ LoggingConfig(name=tag),
117
+ )
118
+
119
+
120
+ if __name__ == "__main__":
121
+ main()
transformer/config.json CHANGED
@@ -1,7 +1,6 @@
1
  {
2
  "_class_name": "QwenImage21Transformer2DModel",
3
- "_diffusers_version": "0.41.0.dev0",
4
- "_name_or_path": "/mnt/hf_cache/hub/models--Qwen--Qwen-Image-2.1/snapshots/b3179ad355be050328e483a9dfdd9e60cd62adfa/transformer",
5
  "attention_head_dim": 128,
6
  "axes_dims_rope": [
7
  16,
@@ -18,43 +17,210 @@
18
  "out_channels": 64,
19
  "patch_size": 1,
20
  "quantization_config": {
21
- "config_groups": {
22
- "group_0": {
23
- "input_activations": {
24
- "dynamic": false,
25
- "num_bits": 4,
26
- "type": "float",
27
- "group_size": 16
28
- },
29
- "weights": {
30
- "dynamic": false,
31
- "num_bits": 4,
32
- "type": "float",
33
- "group_size": 16
34
- },
35
- "targets": [
36
- "Linear"
37
- ]
38
- }
39
- },
40
- "ignore": [
41
- "*img_in",
42
- "*modulation*",
43
- "*norm_out*",
44
- "*proj_out",
45
- "*time_text_embed*",
46
- "*transformer_blocks.0.*",
47
- "*transformer_blocks.1.*",
48
- "*transformer_blocks.30.*",
49
- "*transformer_blocks.31.*",
50
- "*txt_in*"
51
- ],
52
- "quant_algo": "NVFP4",
53
- "producer": {
54
- "name": "modelopt",
55
- "version": "0.46.0rc2"
56
- },
57
- "quant_method": "modelopt",
58
- "quant_type": "nvfp4"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
59
  }
60
- }
 
1
  {
2
  "_class_name": "QwenImage21Transformer2DModel",
3
+ "_diffusers_version": "0.37.0.dev0",
 
4
  "attention_head_dim": 128,
5
  "axes_dims_rope": [
6
  16,
 
17
  "out_channels": 64,
18
  "patch_size": 1,
19
  "quantization_config": {
20
+ "compute_dtype": "bfloat16",
21
+ "quant_method": "nunchaku_lite",
22
+ "svdq_w4a4": {
23
+ "group_size": 16,
24
+ "precision": "nvfp4",
25
+ "rank": 32,
26
+ "targets": [
27
+ "transformer_blocks.2.attn.to_q",
28
+ "transformer_blocks.2.attn.to_k",
29
+ "transformer_blocks.2.attn.to_v",
30
+ "transformer_blocks.2.attn.to_out.0",
31
+ "transformer_blocks.2.img_mlp.proj",
32
+ "transformer_blocks.2.img_mlp.out",
33
+ "transformer_blocks.2.img_mlp.gate_layer",
34
+ "transformer_blocks.3.attn.to_q",
35
+ "transformer_blocks.3.attn.to_k",
36
+ "transformer_blocks.3.attn.to_v",
37
+ "transformer_blocks.3.attn.to_out.0",
38
+ "transformer_blocks.3.img_mlp.proj",
39
+ "transformer_blocks.3.img_mlp.out",
40
+ "transformer_blocks.3.img_mlp.gate_layer",
41
+ "transformer_blocks.4.attn.to_q",
42
+ "transformer_blocks.4.attn.to_k",
43
+ "transformer_blocks.4.attn.to_v",
44
+ "transformer_blocks.4.attn.to_out.0",
45
+ "transformer_blocks.4.img_mlp.proj",
46
+ "transformer_blocks.4.img_mlp.out",
47
+ "transformer_blocks.4.img_mlp.gate_layer",
48
+ "transformer_blocks.5.attn.to_q",
49
+ "transformer_blocks.5.attn.to_k",
50
+ "transformer_blocks.5.attn.to_v",
51
+ "transformer_blocks.5.attn.to_out.0",
52
+ "transformer_blocks.5.img_mlp.proj",
53
+ "transformer_blocks.5.img_mlp.out",
54
+ "transformer_blocks.5.img_mlp.gate_layer",
55
+ "transformer_blocks.6.attn.to_q",
56
+ "transformer_blocks.6.attn.to_k",
57
+ "transformer_blocks.6.attn.to_v",
58
+ "transformer_blocks.6.attn.to_out.0",
59
+ "transformer_blocks.6.img_mlp.proj",
60
+ "transformer_blocks.6.img_mlp.out",
61
+ "transformer_blocks.6.img_mlp.gate_layer",
62
+ "transformer_blocks.7.attn.to_q",
63
+ "transformer_blocks.7.attn.to_k",
64
+ "transformer_blocks.7.attn.to_v",
65
+ "transformer_blocks.7.attn.to_out.0",
66
+ "transformer_blocks.7.img_mlp.proj",
67
+ "transformer_blocks.7.img_mlp.out",
68
+ "transformer_blocks.7.img_mlp.gate_layer",
69
+ "transformer_blocks.8.attn.to_q",
70
+ "transformer_blocks.8.attn.to_k",
71
+ "transformer_blocks.8.attn.to_v",
72
+ "transformer_blocks.8.attn.to_out.0",
73
+ "transformer_blocks.8.img_mlp.proj",
74
+ "transformer_blocks.8.img_mlp.out",
75
+ "transformer_blocks.8.img_mlp.gate_layer",
76
+ "transformer_blocks.9.attn.to_q",
77
+ "transformer_blocks.9.attn.to_k",
78
+ "transformer_blocks.9.attn.to_v",
79
+ "transformer_blocks.9.attn.to_out.0",
80
+ "transformer_blocks.9.img_mlp.proj",
81
+ "transformer_blocks.9.img_mlp.out",
82
+ "transformer_blocks.9.img_mlp.gate_layer",
83
+ "transformer_blocks.10.attn.to_q",
84
+ "transformer_blocks.10.attn.to_k",
85
+ "transformer_blocks.10.attn.to_v",
86
+ "transformer_blocks.10.attn.to_out.0",
87
+ "transformer_blocks.10.img_mlp.proj",
88
+ "transformer_blocks.10.img_mlp.out",
89
+ "transformer_blocks.10.img_mlp.gate_layer",
90
+ "transformer_blocks.11.attn.to_q",
91
+ "transformer_blocks.11.attn.to_k",
92
+ "transformer_blocks.11.attn.to_v",
93
+ "transformer_blocks.11.attn.to_out.0",
94
+ "transformer_blocks.11.img_mlp.proj",
95
+ "transformer_blocks.11.img_mlp.out",
96
+ "transformer_blocks.11.img_mlp.gate_layer",
97
+ "transformer_blocks.12.attn.to_q",
98
+ "transformer_blocks.12.attn.to_k",
99
+ "transformer_blocks.12.attn.to_v",
100
+ "transformer_blocks.12.attn.to_out.0",
101
+ "transformer_blocks.12.img_mlp.proj",
102
+ "transformer_blocks.12.img_mlp.out",
103
+ "transformer_blocks.12.img_mlp.gate_layer",
104
+ "transformer_blocks.13.attn.to_q",
105
+ "transformer_blocks.13.attn.to_k",
106
+ "transformer_blocks.13.attn.to_v",
107
+ "transformer_blocks.13.attn.to_out.0",
108
+ "transformer_blocks.13.img_mlp.proj",
109
+ "transformer_blocks.13.img_mlp.out",
110
+ "transformer_blocks.13.img_mlp.gate_layer",
111
+ "transformer_blocks.14.attn.to_q",
112
+ "transformer_blocks.14.attn.to_k",
113
+ "transformer_blocks.14.attn.to_v",
114
+ "transformer_blocks.14.attn.to_out.0",
115
+ "transformer_blocks.14.img_mlp.proj",
116
+ "transformer_blocks.14.img_mlp.out",
117
+ "transformer_blocks.14.img_mlp.gate_layer",
118
+ "transformer_blocks.15.attn.to_q",
119
+ "transformer_blocks.15.attn.to_k",
120
+ "transformer_blocks.15.attn.to_v",
121
+ "transformer_blocks.15.attn.to_out.0",
122
+ "transformer_blocks.15.img_mlp.proj",
123
+ "transformer_blocks.15.img_mlp.out",
124
+ "transformer_blocks.15.img_mlp.gate_layer",
125
+ "transformer_blocks.16.attn.to_q",
126
+ "transformer_blocks.16.attn.to_k",
127
+ "transformer_blocks.16.attn.to_v",
128
+ "transformer_blocks.16.attn.to_out.0",
129
+ "transformer_blocks.16.img_mlp.proj",
130
+ "transformer_blocks.16.img_mlp.out",
131
+ "transformer_blocks.16.img_mlp.gate_layer",
132
+ "transformer_blocks.17.attn.to_q",
133
+ "transformer_blocks.17.attn.to_k",
134
+ "transformer_blocks.17.attn.to_v",
135
+ "transformer_blocks.17.attn.to_out.0",
136
+ "transformer_blocks.17.img_mlp.proj",
137
+ "transformer_blocks.17.img_mlp.out",
138
+ "transformer_blocks.17.img_mlp.gate_layer",
139
+ "transformer_blocks.18.attn.to_q",
140
+ "transformer_blocks.18.attn.to_k",
141
+ "transformer_blocks.18.attn.to_v",
142
+ "transformer_blocks.18.attn.to_out.0",
143
+ "transformer_blocks.18.img_mlp.proj",
144
+ "transformer_blocks.18.img_mlp.out",
145
+ "transformer_blocks.18.img_mlp.gate_layer",
146
+ "transformer_blocks.19.attn.to_q",
147
+ "transformer_blocks.19.attn.to_k",
148
+ "transformer_blocks.19.attn.to_v",
149
+ "transformer_blocks.19.attn.to_out.0",
150
+ "transformer_blocks.19.img_mlp.proj",
151
+ "transformer_blocks.19.img_mlp.out",
152
+ "transformer_blocks.19.img_mlp.gate_layer",
153
+ "transformer_blocks.20.attn.to_q",
154
+ "transformer_blocks.20.attn.to_k",
155
+ "transformer_blocks.20.attn.to_v",
156
+ "transformer_blocks.20.attn.to_out.0",
157
+ "transformer_blocks.20.img_mlp.proj",
158
+ "transformer_blocks.20.img_mlp.out",
159
+ "transformer_blocks.20.img_mlp.gate_layer",
160
+ "transformer_blocks.21.attn.to_q",
161
+ "transformer_blocks.21.attn.to_k",
162
+ "transformer_blocks.21.attn.to_v",
163
+ "transformer_blocks.21.attn.to_out.0",
164
+ "transformer_blocks.21.img_mlp.proj",
165
+ "transformer_blocks.21.img_mlp.out",
166
+ "transformer_blocks.21.img_mlp.gate_layer",
167
+ "transformer_blocks.22.attn.to_q",
168
+ "transformer_blocks.22.attn.to_k",
169
+ "transformer_blocks.22.attn.to_v",
170
+ "transformer_blocks.22.attn.to_out.0",
171
+ "transformer_blocks.22.img_mlp.proj",
172
+ "transformer_blocks.22.img_mlp.out",
173
+ "transformer_blocks.22.img_mlp.gate_layer",
174
+ "transformer_blocks.23.attn.to_q",
175
+ "transformer_blocks.23.attn.to_k",
176
+ "transformer_blocks.23.attn.to_v",
177
+ "transformer_blocks.23.attn.to_out.0",
178
+ "transformer_blocks.23.img_mlp.proj",
179
+ "transformer_blocks.23.img_mlp.out",
180
+ "transformer_blocks.23.img_mlp.gate_layer",
181
+ "transformer_blocks.24.attn.to_q",
182
+ "transformer_blocks.24.attn.to_k",
183
+ "transformer_blocks.24.attn.to_v",
184
+ "transformer_blocks.24.attn.to_out.0",
185
+ "transformer_blocks.24.img_mlp.proj",
186
+ "transformer_blocks.24.img_mlp.out",
187
+ "transformer_blocks.24.img_mlp.gate_layer",
188
+ "transformer_blocks.25.attn.to_q",
189
+ "transformer_blocks.25.attn.to_k",
190
+ "transformer_blocks.25.attn.to_v",
191
+ "transformer_blocks.25.attn.to_out.0",
192
+ "transformer_blocks.25.img_mlp.proj",
193
+ "transformer_blocks.25.img_mlp.out",
194
+ "transformer_blocks.25.img_mlp.gate_layer",
195
+ "transformer_blocks.26.attn.to_q",
196
+ "transformer_blocks.26.attn.to_k",
197
+ "transformer_blocks.26.attn.to_v",
198
+ "transformer_blocks.26.attn.to_out.0",
199
+ "transformer_blocks.26.img_mlp.proj",
200
+ "transformer_blocks.26.img_mlp.out",
201
+ "transformer_blocks.26.img_mlp.gate_layer",
202
+ "transformer_blocks.27.attn.to_q",
203
+ "transformer_blocks.27.attn.to_k",
204
+ "transformer_blocks.27.attn.to_v",
205
+ "transformer_blocks.27.attn.to_out.0",
206
+ "transformer_blocks.27.img_mlp.proj",
207
+ "transformer_blocks.27.img_mlp.out",
208
+ "transformer_blocks.27.img_mlp.gate_layer",
209
+ "transformer_blocks.28.attn.to_q",
210
+ "transformer_blocks.28.attn.to_k",
211
+ "transformer_blocks.28.attn.to_v",
212
+ "transformer_blocks.28.attn.to_out.0",
213
+ "transformer_blocks.28.img_mlp.proj",
214
+ "transformer_blocks.28.img_mlp.out",
215
+ "transformer_blocks.28.img_mlp.gate_layer",
216
+ "transformer_blocks.29.attn.to_q",
217
+ "transformer_blocks.29.attn.to_k",
218
+ "transformer_blocks.29.attn.to_v",
219
+ "transformer_blocks.29.attn.to_out.0",
220
+ "transformer_blocks.29.img_mlp.proj",
221
+ "transformer_blocks.29.img_mlp.out",
222
+ "transformer_blocks.29.img_mlp.gate_layer"
223
+ ]
224
+ }
225
  }
226
+ }
transformer/diffusion_pytorch_model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a6c3080b764344729ef95846faa3a0a01eb36bdef93f785ed07f4ac5751ad76a
3
- size 5451672424
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0304fc3e84b13663b439601380066754ef9b92025e159275abe163bc5f1ceeb4
3
+ size 5603299808
transformer/modeling_nunchaku_qwenimage21.py ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Qwen-Image-2.1 transformer, unchanged, loaded as a custom component.
2
+
3
+ Its only job is importing `nunchaku_kernels`, so the NVFP4 kernels are set up even when
4
+ Diffusers loads this component before the text encoder. The Nunchaku Lite quantization
5
+ itself is handled by Diffusers' built-in quantizer from `config.json`.
6
+ """
7
+
8
+ from diffusers import QwenImage21Transformer2DModel
9
+
10
+ from .nunchaku_kernels import KERNELS # noqa: F401
11
+
12
+
13
+ class NunchakuQwenImage21Transformer2DModel(QwenImage21Transformer2DModel):
14
+ pass
transformer/nunchaku_kernels.py ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Shared by text_encoder/ and transformer/: whichever component Diffusers loads first sets this up.
2
+ #
3
+ # Diffusers fetches the Nunchaku Lite kernels from `rootonchair/nunchaku-lite-kernels`, which is
4
+ # no longer downloadable. Loading this repo with trust_remote_code=True already runs this code,
5
+ # so point that kernel name at the rebuilt copy before Diffusers imports its Nunchaku utilities.
6
+ import os
7
+
8
+ from huggingface_hub import snapshot_download
9
+
10
+ KERNELS = "rootonchair/nunchaku-lite-kernels"
11
+ if KERNELS not in os.environ.get("LOCAL_KERNELS", ""):
12
+ local = f"{KERNELS}={snapshot_download('joseplcam/nunchaku-lite-kernels')}"
13
+ os.environ["LOCAL_KERNELS"] = ":".join(filter(None, [os.environ.get("LOCAL_KERNELS"), local]))
14
+ if "DIFFUSERS_TRUST_REMOTE_KERNELS" not in os.environ:
15
+ # Diffusers read this variable into a constant at import time, so set both.
16
+ import diffusers.utils.constants
17
+
18
+ os.environ["DIFFUSERS_TRUST_REMOTE_KERNELS"] = "true"
19
+ diffusers.utils.constants.DIFFUSERS_TRUST_REMOTE_KERNELS = True
vae/config.json CHANGED
@@ -1,6 +1,6 @@
1
  {
2
  "_class_name": "AutoencoderKLQwenImage21",
3
- "_diffusers_version": "0.37.0.dev0",
4
  "attn_scales": [],
5
  "base_dim": 96,
6
  "decoder_base_dim": 144,
 
1
  {
2
  "_class_name": "AutoencoderKLQwenImage21",
3
+ "_diffusers_version": "0.41.0.dev0",
4
  "attn_scales": [],
5
  "base_dim": 96,
6
  "decoder_base_dim": 144,
vae/diffusion_pytorch_model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a07a1b7c4ee2966a1b3bdc37de9b4f983d56937e46619f709a80b6e490675417
3
- size 1350989512
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:71879ffd5321e6d10c3c87513e2b474b1252efa7f3dec2969214a9bf06a6dd5c
3
+ size 675508656