NidAll's picture
Update README.md
fa3fb50 verified
|
Raw
History Blame Contribute Delete
3.31 kB
---
license: other
base_model: Qwen/Qwen-Image-2.1
base_model_relation: quantized
pipeline_tag: text-to-image
library_name: diffusion-single-file
tags:
- qwen
- qwen-image
- qwen-image-2.1
- comfyui
- quantization
- mixed-precision
- w4a8
- int8
- convrot
- safetensors
- image-generation
- image-editing
---
# Qwen-Image 2.1 Mixed Balanced
A **ComfyUI-native mixed-precision quantization** of [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1), created with [NidAll/comfyui-native-quantizer](https://github.com/NidAll/comfyui-native-quantizer).
The goal is simple: keep Qwen-Image 2.1's image quality and typography while making the diffusion transformer much smaller and friendlier to low-VRAM GPUs.
## Size
| Checkpoint | Size |
| --- | ---: |
| BF16 source | **13.25 GiB** |
| Mixed Balanced | **3.91 GiB** |
| Reduction | **70.5%** |
The conversion quantized **192 transformer linears**, kept **73 tensors** at source precision, and passed **65 validation checks with 0 warnings and 0 failures**.
## Quantization
The checkpoint was generated with:
```bash
comfyquant-cuda qwen_image_2.1_bf16.safetensors \
--format mixed \
--profile balanced \
--experimental \
--target-runtime cuda \
--output qwen_image_2.1_mixed_balanced.safetensors
```
Large attention and MLP linears are quantized, while sensitive paths such as inputs, timestep conditioning, modulation, attention norms, final normalization, and output projection remain at source precision.
Conversion summary:
```text
architecture : qwen_image21
format : mixed / balanced
quantized : 192 layers
kept : 73 tensors
mean error : 0.0731
compression : 0.563 bytes/param
validation : 65 passed, 0 warnings, 0 failed
```
## Example generations
### Painterly composition
![Black hole example](examples/Qwen_image_2.1_00033.png)
### Typography-heavy comic
![Linux comic example](examples/Qwen_image_2.1_00035.png)
The comic is a useful stress test because it combines long text, multiple panels, small labels, repeated characters, and dense composition.
## ComfyUI
Place the checkpoint in:
```text
ComfyUI/models/diffusion_models/
```
Then use it with a normal Qwen-Image 2.1 workflow.
Recommended companion components for low-VRAM setups:
- **Qwen3-VL 8B W4A8** text encoder
- **Qwen-Image 2.1 BF16 VAE**
The text encoder and VAE are not included in this repository.
## Quality notes
Initial testing shows very strong retention compared with the official INT8 ConvRot model, including typography-heavy generations.
Outputs are not expected to be pixel-identical to BF16 or INT8 at the same seed. Small numerical differences can change the diffusion trajectory while preserving overall prompt adherence and visual quality.
This checkpoint is experimental, so BF16 or the official INT8 model remain useful references for maximum reproducibility.
## Credits
- [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1)
- [Comfy-Org/Qwen-Image-2.1](https://huggingface.co/Comfy-Org/Qwen-Image-2.1)
- [NidAll/comfyui-native-quantizer](https://github.com/NidAll/comfyui-native-quantizer)
## License
This is an independent quantized derivative of Qwen-Image 2.1. Use remains subject to the applicable upstream Qwen-Image 2.1 license terms.