Update README.md
Browse files
README.md
CHANGED
|
@@ -18,6 +18,15 @@ base_model: Qwen/Qwen-Image-2.1
|
|
| 18 |
base_model_relation: quantized
|
| 19 |
---
|
| 20 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 21 |
# Qwen-Image-2.1 7B — INT4 (W4A8) & INT6 (W6A8) ConvRot for ComfyUI
|
| 22 |
|
| 23 |
INT4 and INT6 quantized weights of **Qwen-Image-2.1** for fast, low-VRAM inference in ComfyUI.
|
|
|
|
| 18 |
base_model_relation: quantized
|
| 19 |
---
|
| 20 |
|
| 21 |
+
## About these variants
|
| 22 |
+
|
| 23 |
+
Four precision levels are available for Qwen-Image-2.1, trading VRAM and speed against generation quality:
|
| 24 |
+
|
| 25 |
+
- **BF16** — full-precision reference (~14 GB). Highest quality, largest footprint.
|
| 26 |
+
- **INT8 (W8A8)** — 8-bit weights + 8-bit activations (~7 GB). Near-lossless quality, ~2× smaller than BF16, runs on INT8 tensor cores (RTX 30-series and up).
|
| 27 |
+
- **INT6 (W6A8)** — 6-bit weights + 8-bit activations (~5.5 GB). Middle-ground: smaller than INT8 with only a small quality trade-off.
|
| 28 |
+
- **INT4 (W4A8)** — 4-bit weights + 8-bit activations (~4 GB). Smallest footprint; uses ConvRot with a per-tensor codebook that decodes to INT8 for compute, so it runs on the same INT8 hardware as W8A8. Larger quality trade-off than INT6.
|
| 29 |
+
|
| 30 |
# Qwen-Image-2.1 7B — INT4 (W4A8) & INT6 (W6A8) ConvRot for ComfyUI
|
| 31 |
|
| 32 |
INT4 and INT6 quantized weights of **Qwen-Image-2.1** for fast, low-VRAM inference in ComfyUI.
|