tsolful's picture
Update README.md
31d2357 verified
|
Raw History Blame Contribute Delete
3.19 kB
---
license: other
license_name: qwen-research
license_link: https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE
pipeline_tag: text-to-image
tags:
- comfyui
- qwen-image-2.1
- text-to-image
- image-to-image
- image-generation
- quantized
- qwen
- image-generation
- image-editing
- rgba
base_model: Qwen/Qwen-Image-2.1
base_model_relation: quantized
---
## About these variants
Four precision levels are available for Qwen-Image-2.1, trading VRAM and speed against generation quality:
- **BF16** β€” full-precision reference (~14 GB). Highest quality, largest footprint.
- **INT8 (W8A8)** β€” 8-bit weights + 8-bit activations (~7 GB). Near-lossless quality, ~2Γ— smaller than BF16, runs on INT8 tensor cores (RTX 30-series and up).
- **INT6 (W6A8)** β€” 6-bit weights + 8-bit activations (~5.5 GB). Middle-ground: smaller than INT8 with only a small quality trade-off.
- **INT4 (W4A8)** β€” 4-bit weights + 8-bit activations (~4 GB). Smallest footprint; uses ConvRot with a per-tensor codebook that decodes to INT8 for compute, so it runs on the same INT8 hardware as W8A8. Larger quality trade-off than INT6.
# Qwen-Image-2.1 7B β€” INT4 (W4A8) & INT6 (W6A8) ConvRot for ComfyUI
INT4 and INT6 quantized weights of **Qwen-Image-2.1** for fast, low-VRAM inference in ComfyUI.
This is a modified (quantized) version of the Qwen-Image-2.1 model. It is not an official Qwen release and is not endorsed by the Qwen team.
## About Qwen-Image-2.1
A unified text-to-image generation and image editing model in the Qwen family. With just 7B parameters in its visual generation component (32 Single-Stream DiT layers), Qwen-Image-2.1 balances generation quality, inference efficiency, and versatility.
### Highlights
- **Efficient Image Generation** β€” combines strong visual performance with fast inference and a compact design, making high-quality image creation accessible across a wide range of creative workflows.
- **Flexible Creative Control** β€” supports diverse inputs, outputs, and localized edits, giving creators the flexibility to explore ideas and refine details within a unified workflow.
### Key improvements in 2.1
- **Compact and Efficient** β€” lightweight architecture with mixed-granularity attention and prefix KV cache reuse delivers strong image quality at low computational cost.
- **Native Transparency, Unified Creation and Editing** β€” generate regular or transparent (RGBA) images from text, edit transparent layers, and extract subjects from photographs β€” all in one model.
- **Versatile Editing** β€” support up to 10 reference images, specify local edits via circles, painted annotations, or separate masks, and preserve identity for people and products.
- **Realistic Textures and Refined Aesthetics** β€” improved typography, portrait lighting, and fine details for more visually compelling results.
## License
Qwen-Image-2.1 is licensed under the [Qwen Research License](https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE). These files are a quantized derivative and are distributed under the same license β€” see the original repository for full terms, permitted uses, and any commercial-use restrictions.