tsolful commited on
Commit
31d2357
·
verified ·
1 Parent(s): f4b000a

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +9 -0
README.md CHANGED
@@ -18,6 +18,15 @@ base_model: Qwen/Qwen-Image-2.1
18
  base_model_relation: quantized
19
  ---
20
 
 
 
 
 
 
 
 
 
 
21
  # Qwen-Image-2.1 7B — INT4 (W4A8) & INT6 (W6A8) ConvRot for ComfyUI
22
 
23
  INT4 and INT6 quantized weights of **Qwen-Image-2.1** for fast, low-VRAM inference in ComfyUI.
 
18
  base_model_relation: quantized
19
  ---
20
 
21
+ ## About these variants
22
+
23
+ Four precision levels are available for Qwen-Image-2.1, trading VRAM and speed against generation quality:
24
+
25
+ - **BF16** — full-precision reference (~14 GB). Highest quality, largest footprint.
26
+ - **INT8 (W8A8)** — 8-bit weights + 8-bit activations (~7 GB). Near-lossless quality, ~2× smaller than BF16, runs on INT8 tensor cores (RTX 30-series and up).
27
+ - **INT6 (W6A8)** — 6-bit weights + 8-bit activations (~5.5 GB). Middle-ground: smaller than INT8 with only a small quality trade-off.
28
+ - **INT4 (W4A8)** — 4-bit weights + 8-bit activations (~4 GB). Smallest footprint; uses ConvRot with a per-tensor codebook that decodes to INT8 for compute, so it runs on the same INT8 hardware as W8A8. Larger quality trade-off than INT6.
29
+
30
  # Qwen-Image-2.1 7B — INT4 (W4A8) & INT6 (W6A8) ConvRot for ComfyUI
31
 
32
  INT4 and INT6 quantized weights of **Qwen-Image-2.1** for fast, low-VRAM inference in ComfyUI.