File size: 6,624 Bytes
2e77d03
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
---
license: other
license_name: qwen-research-license
license_link: LICENSE
base_model: Comfy-Org/Qwen-Image-2.1
base_model_relation: quantized
pipeline_tag: text-to-image
tags:
- qwen-image
- qwen-image-2.1
- nvfp4
- comfyui
- diffusion-single-file
- quantized
- image-editing
- rtx-5090
---

# πŸš€ Prism Image 2.1 β€” NVFP4 for ComfyUI

**The image model AND the text encoder. 12.42 GB for the complete package.**

I converted Qwen Image 2.1's image transformer and Qwen3-VL 8B language encoder
directly from the official BF16 weights into native ComfyUI NVFP4. The vision
tower, embeddings, output head, critical image-model tensors, and VAE stay BF16.
Real Blackwell FP4 kernels in both denoising and text conditioning, with an
included small runtime patch for the latter. No new nodes.

**Built with Qwen.** Research/evaluation use only under the included license.

**RTX 5090: 6.746 s median for a complete 1024Γ—1024, 40-step run** versus
15.629 s with BF16 and 7.553 s with official INT8. Three warm runs per variant,
every node executed, including text encoding. This one-prompt benchmark used the
encoder patch; [full method and limits](VALIDATION.md).

![Reference photo and the NVFP4 product edit](assets/editing.jpg)

## πŸ“¦ Download once, wire nothing

Tested with **ComfyUI 0.36.0**, PyTorch **2.14.0+cu130**, comfy-kitchen **0.2.35**,
and an **RTX 5090**. Use a ComfyUI build with `TextEncodeQwenImage21` and native
NVFP4 support. Older builds without the Qwen Image 2.1 integration will not work.

| File | Where it goes | Size |
|---|---|---:|
| [prism_image_2.1_nvfp4.safetensors](diffusion_models/prism_image_2.1_nvfp4.safetensors) | `ComfyUI/models/diffusion_models/` | 4.20 GB |
| [prism_qwen3vl_8b_nvfp4.safetensors](text_encoders/prism_qwen3vl_8b_nvfp4.safetensors) | `ComfyUI/models/text_encoders/` | 7.55 GB |
| [qwen_image_2.1_vae_bf16.safetensors](vae/qwen_image_2.1_vae_bf16.safetensors) | `ComfyUI/models/vae/` | 0.68 GB |

The original BF16 set is 32.44 GB. This package is **61.7% smaller on disk**.
File sizes are decimal GB, not a VRAM requirement. Keep your existing identical
VAE if you already have it.

1. Put the three files in the folders above and refresh ComfyUI's model list.
2. For accelerated NVFP4 text conditioning, apply the [runtime patch](runtime/README.md)
   and restart. Without it, the same models and nodes work, but text conditioning
   uses dequantized weights and FP32 activations. Denoising still uses NVFP4.
3. Import one of the workflows below using **Ctrl+O**, then click **Run**.
4. For editing, use **Upload** in Load Image to select the included
   [reference image](input/prism_edit_reference.png), or your own image.

| Working workflow | Purpose |
|---|---|
| [01 β€” Text to image](workflows/01_Text_to_Image.json) | 1024Γ—1024 product example; change the prompt and latent size |
| [02 β€” Image editing](workflows/02_Image_Editing.json) | One reference, coherent material/color/object changes |
| [03 β€” Transparent RGBA](workflows/03_Transparent_RGBA.json) | Native alpha output; no background-removal node |
| [04 β€” 2K typography](workflows/04_2K_Typography.json) | 2048Γ—2048 poster with exact requested lettering |

The workflows use core `UNETLoader`, `CLIPLoader` (type **qwen_image**),
`TextEncodeQwenImage21`, `KSampler`, and VAE nodes. Leave diffusion weight dtype
at **default**. Defaults are 40 steps, Euler, simple scheduler, CFG 1, denoise 1.
The encoder node's resolution controls reference resizing; text-to-image output
size comes from Empty Latent Image. API equivalents are in [workflows/api](workflows/api).

## πŸ‘€ Same prompts, same seeds

![BF16 and NVFP4 portrait comparison](assets/ab_portrait_20260920.jpg)

![BF16 and NVFP4 typography comparison](assets/ab_typography_20260921.jpg)

All **24 matched pairs** are in [COMPARISONS.md](COMPARISONS.md). Download
`gallery.html` and its `assets` folder together for the offline comparison slider.
See [PROMPTING.md](PROMPTING.md) for text, editing, lettering and transparency tips.

See [VALIDATION.md](VALIDATION.md) for the test set, measured results, precision
tradeoffs, and execution requirements. The comparisons use BF16 source weights
against this complete accelerated NVFP4 package. Quantization changes results;
this is not a promise of pixel-identical output or a universal quality ranking.

## 🧠 What's actually quantized?

| Component | Policy |
|---|---|
| Image transformer | 192 attention Q/K/V/output and fused MLP gate/up/output matrices: NVFP4 |
| Qwen3-VL language backbone | 252 attention and MLP matrices: NVFP4 |
| Vision tower, token embeddings, language head | Original BF16 tensors retained |
| Image-model input/output, timestep, modulation, normalization | Original BF16 tensors retained |
| VAE | Original BF16 file, unchanged |
| Scales | E4M3 FP8 blocks of 16, FP32 tensor scales |

Native packed E2M1 FP4, native scale layout, and per-layer `comfy_quant` metadata.
This is a mixed-precision NVFP4 stack, not every parameter forced into four bits.
No GGUF loader, no requantization of the INT8 release, no dropped vision tower.
No calibration or fine-tuning was used for this conversion.

## πŸ”§ Reproduce it

Source repository: [Comfy-Org/Qwen-Image-2.1](https://huggingface.co/Comfy-Org/Qwen-Image-2.1),
revision `ace0edeb3791a594ddfa36ed5f41a178a394e921`.

Use the Python environment of the tested ComfyUI installation:

```powershell
python tools/convert.py path/to/qwen_image_2.1_bf16.safetensors diffusion_models/prism_image_2.1_nvfp4.safetensors --component dit --comfy-root path/to/ComfyUI
python tools/convert.py path/to/qwen3vl_8b_bf16.safetensors text_encoders/prism_qwen3vl_8b_nvfp4.safetensors --component encoder --comfy-root path/to/ComfyUI
```

The converter uses installed Comfy Kitchen operations, validates matrix alignment,
and emits tensor manifests with source/output hashes and reconstruction error.
`tools/verify.py SOURCE CONVERTED` checks the stored format, checkpoint hash, and
exact equality of every protected tensor. Final file checksums are in `SHA256SUMS`.

## βš–οΈ License & credits

Qwen created the base model; Comfy-Org provided the native BF16 checkpoint package;
ComfyUI and Comfy Kitchen provide loading and GPU execution. Independent conversion,
workflow packaging, and RTX 5090 validation by **BennyDaBall_OG**.

The [Qwen Research License](LICENSE) limits use to non-commercial research and
evaluation; commercial use requires a separate license from Qwen. Read the full
license and [NOTICE](NOTICE). The included ComfyUI patch remains GPL-3.0.
This repo is not affiliated with Qwen or Comfy-Org.