---
license: other
license_name: qwen-research
license_link: LICENSE
base_model:
- Qwen/Qwen-Image-2.1
tags:
- quantized
- nvfp4
- comfyui
- qwen-image
- text-to-image
---
# Qwen-Image-2.1 DiT — NVFP4 (ComfyUI)
> ⚠️ **ComfyUI format only.** These files use ComfyUI's native quantized-checkpoint
> layout (`comfy_quant` tensors + `_quantization_metadata`). They will **not** load
> in `diffusers`, and they need a ComfyUI build with `QwenImage21` support
> (merged after 2026-09-14). If you want the model for `diffusers`, use the
> [official bf16 release](https://huggingface.co/Qwen/Qwen-Image-2.1).
> **Not affiliated with, or endorsed by, Alibaba / Qwen.** Community derivative of
> [`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1), redistributed
> under the **Qwen Research License** (copy included as `LICENSE`, per §3.a).
> **Non-commercial use only** (§1.i / §2.a); commercial use needs a separate licence
> from Qwen (`model-business@notice.qwencloud.com`, §2.b).
NVFP4 quantizations of the **diffusion transformer**. Comfy-Org ships `bf16` and
`int8_convrot`; **there is no official NVFP4 build**, so these are new.
---
## Showcase
Every image below was generated with **these NVFP4 files** (the default `qwen_image_2.1_nvfp4.safetensors`) driven by the **PE-T2I rewriter** — each from a single short prompt plus a photographic style, through a small ComfyUI node: **[ComfyUI-QwenImage-PhotoStyles](https://github.com/pottokao-dotcom/ComfyUI-QwenImage-PhotoStyles)**. Curious how it's put together? The node and its full recipe are in that repo.
 Black Fury |
 Geometry of Light |
 Warm Documentary |
 White Leap |
 Cold Power |
 Hyper Chic East |
 Private Diary |
 Color Behind Glass |
 Frozen Film Still |
 Fairytale Giants |
 Vivid Garden |
 Playful Film |
---
## Pick a file
| File | Size | Start here if… |
|---|---:|---|
| **`qwen_image_2.1_nvfp4.safetensors`** | **3.91 GB** | **you just want the smallest thing that works — this is the default** |
| `qwen_image_2.1_nvfp4_T2.safetensors` | 4.29 GB | you see artefacts on hard material and want a little more headroom |
| `qwen_image_2.1_nvfp4_T3.safetensors` | 4.67 GB | you want the most conservative option in the set |
Source is 13.25 GB bf16, so these are **29.5 % / 32.4 % / 35.2 %** of the original.
### What actually differs between them
The model has **32 transformer blocks**, each with 6 quantizable layers
(`attn.to_q/k/v/out.0`, `img_mlp.gate_up`, `img_mlp.out`) = 192 layers total.
The only axis that changes between tiers is **how many blocks at the head and tail
are kept at 8-bit instead of 4-bit**:
| | 4-bit NVFP4 (group 16) | 8-bit MXFP8 (group 32) | bf16 |
|---|---|---|---|
| **T1** (3.91 GB) | all 192 layers | — | 73 tensors |
| **T2** (4.29 GB) | 168 layers | 24 layers — blocks **0, 1, 30, 31** | 73 tensors |
| **T3** (4.67 GB) | 144 layers | 48 layers — blocks **0–3, 28–31** | 73 tensors |
The bf16 protection set is **identical in all three** — the tiers differ only in
that 4-bit ↔ 8-bit trade at the ends of the stack.
### Does the extra size buy anything?
On a like-for-like test — same seed, same prompt, same text encoder — the three
tiers produced near-identical images: **mean per-pixel difference of 3–4 out of
255**. The extra 0.4–0.8 GB did not show a visible benefit there.
That was one prompt and one seed, so treat it as a starting point rather than a
verdict. Faces, small text and dense texture are the places where a difference
would show up first, so if you work on those, T2/T3 are worth a try.
---
## What is protected, and why
| Layers | Count | Precision |
|---|---:|---|
| `attn.to_q/k/v/out`, `img_mlp.gate_up/out` | 192 | NVFP4, or MXFP8 in T2/T3 |
| `img_in`, `txt_in.in_layer/out_layer` | 3 | **bf16** |
| `time_text_embed.timestep_embedder.linear_1/2` | 2 | **bf16** |
| `modulation.1` | 1 | **bf16** |
| `norm_out.linear`, `proj_out` | 2 | **bf16** |
| norms / biases (non-2-D) | 65 | bf16 |
98.1 % of parameters are quantized; the protected 1.9 % costs only ~240 MB.
**Why the timestep path stays bf16.** A DiT's `time_text_embed` and modulation decide
*how far each denoising step moves*. Measurements on a distilled model in this family
showed that when timestep conditioning is rewritten, the model stops working outside
the exact schedule it was trained for. Keeping that path in bf16 costs almost nothing
and removes a whole class of risk.
Every 2-D weight in this model has `in_features` divisible by 16, so **no padding is
needed anywhere** — the quantization is clean.
## Audit
| Check | Result |
|---|---|
| Metadata | `format_version 1.0`, 192 layers declared |
| Tensor bookkeeping | 0 missing, 0 extra; `comfy_quant` matches metadata for every layer |
| NVFP4 round-trip error | median **9.44 %**, max 9.49 % (24-layer sample) |
| MXFP8 round-trip error | median **2.65 %** — 3.6× tighter than NVFP4, as the bit budget predicts |
| Protected layers | 40/40 sampled are **byte-identical** to the bf16 source |
| Loads in ComfyUI | ✅ all three, detected as `QwenImage21` |
| End-to-end | ✅ 1024×1024, 25 steps, **~20–22 s** on a GB10 |
The MXFP8 number matters: had the 8-bit path been written wrong, it would have landed
at or above the 4-bit error. It didn't.
## Usage
Place in `ComfyUI/models/diffusion_models/`, then use the official
[Qwen-Image-2.1 text-to-image workflow](https://github.com/Comfy-Org/workflow_templates/blob/main/templates/image_qwen_image_2_1_t2i.json):
```
UNETLoader → (this file)
CLIPLoader → qwen3vl_8b_*, type: qwen_image
TextEncodeQwenImage21
VAELoader → qwen_image_2.1_vae_bf16.safetensors
KSampler → 25 steps, cfg 1.0, euler, simple
```
Older ComfyUI builds report `UNSUPPORTED DIFFUSION MODEL` for Qwen-Image-2.1 in
general — **including the official bf16 files** — so if you see that, update ComfyUI
rather than suspecting these weights.
## Getting legible text

*Dense drink labels stay legible while the bottles keep their glass/liquid translucency — the cfg-schedule handoff below.*
The `cfg 1.0` above is the fast default, but this DiT is the **full model, not a
distilled few-step one** — so classifier-free guidance is a real lever, and text is
where it shows most:
- **cfg 1.0** — best photorealism and material feel (glass/liquid translucency, mood),
but dense small text comes out as plausible-looking gibberish.
- **cfg ~3.0 + a negative prompt** — text snaps to legible (brand names,
prices, whole menu rows), at the cost of flatter, more "plastic" materials and higher
saturation. Put `oversaturated, overexposed` in the negative so it doesn't scorch.
- **cfg ~2.0** — a reasonable single-pass middle ground.

*Same bottle crop: cfg 1.0 keeps the glass/liquid translucency but the labels are gibberish; cfg 3.0 makes labels legible but the bottles go flat and plastic; the schedule handoff keeps most of the translucency while the labels turn readable.*
No single cfg is best at both a great photo *and* great text, so the best result is a
**schedule handoff**: run the first ~2/3 of the steps at cfg 1.0 (structure and
materials lock in), then the last ~1/3 at cfg 3.0 (text sharpens without
re-plasticising the materials, which are already set). One denoising trajectory, one
seed — two `KSamplerAdvanced` nodes:
```
KSamplerAdvanced #1 add_noise enable, cfg 1.0, start_at_step 0, end_at_step 17, return_with_leftover_noise enable
KSamplerAdvanced #2 add_noise disable, cfg 3.0, start_at_step 17, end_at_step 25, return_with_leftover_noise disable
```
Both take the same model / positive / negative; #2's `latent_image` is #1's output.
Move the split later for more material fidelity, earlier for sharper text.
### Speed
`cfg > 1` evaluates the negative branch too, so cfg-1 steps cost 1× and cfg-3 steps 2×.
The handoff only pays that on its final third. GB10, 1024×1280, 25 steps, T3:
| Setting | Forward passes | Time |
|---|---:|---:|
| cfg 1.0 (all steps) | 25 | **27 s** |
| **handoff (2/3 cfg1 → 1/3 cfg3)** | 33 | **33 s** |
| cfg 3.0 (all steps) | 50 | 48 s |
The handoff buys cfg-3 text quality for ~20 % over the cfg-1 baseline, and runs a third
faster than cfg 3.0 the whole way.
**Legible signage also needs the text spelled out.** The model only renders text it is
*given* — put exact strings in the prompt (`a sign that reads "居酒屋"`); it invents
gibberish for any sign you merely describe. The "rewriter" is a prompt-enhancement LLM — Qwen's own **PE-T2I** (*Prompt Enhancement for text-to-image*), which expands a short request into a detailed prompt; a refusal-ablated build is [`…-PE-T2I-Heretic`](https://huggingface.co/pottokao/Qwen-Image-2.1-PE-T2I-Heretic). It runs *before* the text encoder, and is not the text encoder itself. A prompt-rewriter that lays signs out
spatially (distinct positions, generous size) lets it render many at once — the
difference between readable and "pretty neon soup".

*Five scenes, each row: **left** = a terse prompt straight to the model, **right** = the same request expanded by a prompt-rewriter. Same DiT, same seed. Without it you get atmospheric but garbled neon; with it, legible signage (`歌舞伎町`, `FamilyMart 120種類`, a full ramen menu, `自動販売機`) and a composed frame. The rewriter does the interpretive heavy lifting.*
## Pitfall worth knowing (T2/T3)
`TensorCoreMXFP8Layout.quantize()` returns its scale as `float8_e8m0`, which
ComfyUI's safetensors loader cannot parse — `KeyError: 'F8_E8M0'`, and the file simply
refuses to load. The official checkpoints store that bit pattern as `uint8` and
reinterpret it at load time; these files do the same.
The all-NVFP4 tier never touches that code path, so **a bug in the MXFP8 route stays
invisible until you test T2/T3 specifically** — verifying T1 and assuming the rest
follow is exactly how it gets missed.
## Text encoder
The other half of the pipeline, with refusal behaviour ablated
(5/100 refusals @ KL 0.0220):
| Repo | Format | Runtime |
|---|---|---|
| [`…-Text-Encoder-Heretic`](https://huggingface.co/pottokao/Qwen-Image-2.1-Text-Encoder-Heretic) | bf16, 16.33 GB | `transformers` / any |
| [`…-Heretic-NVFP4`](https://huggingface.co/pottokao/Qwen-Image-2.1-Text-Encoder-Heretic-NVFP4) | NVFP4, 5.87 GB | **ComfyUI** |
| [`…-Heretic-W4A8`](https://huggingface.co/pottokao/Qwen-Image-2.1-Text-Encoder-Heretic-W4A8) | W4A8 INT8, 5.88 GB | **ComfyUI** |
| [`…-Heretic-GGUF`](https://huggingface.co/pottokao/Qwen-Image-2.1-Text-Encoder-Heretic-GGUF) | GGUF Q4_K_M + mmproj, 5.76 GB | **llama.cpp / Mac** |
This DiT + the NVFP4 text encoder is a complete **~9.8 GB** pipeline.