File size: 12,343 Bytes
4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d 4825e8b cbe716d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 | ---
license: other
license_name: qwen-research
license_link: https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE
library_name: optimum-intel
pipeline_tag: text-to-image
base_model: Qwen/Qwen-Image-2.1
tags:
- openvino
- optimum-intel
- int4
- nncf
- text-to-image
---
# Qwen-Image-2.1 — OpenVINO INT4
將 [`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1) 轉換為 **OpenVINO INT4** 權重,並以 **純 CPU** 實測出圖。
`optimum-intel` 匯出 → NNCF weight-only INT4 量化(耗時 115 s)→ **FP16 43.6 GB → INT4 13.12 GB(↓ ~70%,約 3.3×)**。
> ### ⚠️ 環境需求與其他 repo 不同
>
> `QwenImage21Pipeline` 在 diffusers **0.40.0 之後**才加入,因此 `diffusers==0.37.1` **無法使用**。
> 本 repo 需要 diffusers 主分支版本,且需對 `transformers` 與 `optimum-intel` 套用兩處小幅度修改
>(詳見下方安裝說明與 `REPORT.md`)。
---
## 目錄
- [快速資訊](#快速資訊)
- [特色](#特色)
- [安裝](#安裝)
- [快速開始](#快速開始)
- [推論參數建議](#推論參數建議)
- [範例結果](#範例結果)
- [效能實測摘要](#效能實測摘要)
- [檔案結構](#檔案結構)
- [從零復現](#從零復現)
- [已知限制](#已知限制)
- [授權與出處](#授權與出處)
---
## 快速資訊
| 項目 | 內容 |
| --- | --- |
| **基座模型** | [`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1) |
| **Pipeline** | `QwenImage21Pipeline` |
| **Scheduler** | `FlowMatchEulerDiscreteScheduler`(`shift=1.0`、`shift_terminal=0.02`、`use_dynamic_shifting=true`) |
| **量化格式** | `transformer`、`text_encoder`、`text_encoder_i2i` → INT4;其餘 → INT8 |
| **模型大小** | FP16 43.6 GB → **INT4 13.12 GB**(↓ ~70%,約 3.3×) |
| **推論裝置** | CPU(OpenVINO CPU plugin,無需 GPU) |
| **推薦參數** | `num_inference_steps=40`、`true_cfg_scale=1.0`、`1024×1024` |
| **轉換工具** | optimum-intel 2.2.x / optimum 2.3.0 / OpenVINO 2026.4.0 / NNCF 3.4.0 |
## 特色
- **三個元件 INT4**:Transformer 與兩個 Qwen3-VL text encoder 全部 INT4,`text_encoder_i2i` 是編輯功能用的第二個 text encoder。
- **⚠️ 需要 diffusers 主分支**:`QwenImage21Pipeline` 需要 `diffusers==0.41.0.dev0`與 `transformers==5.10.4`,與同批其他 repo 的版本需求不同。
- **體積 ↓ 70%**:FP16 43.6 GB → INT4 13.12 GB(約 3.3×)。
- **支援編輯能力**:保留了 image-to-image 的 `text_encoder_i2i` 與 `vision_encoder`。
- **⚠️ 速度最慢**:1024×1024 / 40 steps 平均 870 s / 張(約 14.5 分鐘)。
- **記憶體需求最高**:行程結束時 RSS 達 40.2 GB。
## 安裝
```bash
pip install diffusers==0.41.0.dev0 transformers==5.10.4 tokenizers==0.22.2 huggingface-hub==1.33.0 optimum==2.3.0 optimum-intel==2.3.0.dev0+32a317a openvino==2026.4.0 nncf==3.4.0 torch==2.14.0 pillow==12.3.0 psutil==7.2.2
```
`nncf` 僅重新量化時需要。
## 快速開始
```python
import torch
from optimum.intel import OVDiffusionPipeline
pipe = OVDiffusionPipeline.from_pretrained("HelloSun/Qwen-Image-2.1-OpenVINO-INT4", compile=True)
image = pipe(
prompt="Astronaut in a jungle, cold color palette, muted colors, "
"detailed, 8k, photorealistic, cinematic lighting",
height=1024, width=1024,
num_inference_steps=40,
true_cfg_scale=1.0,
generator=torch.Generator().manual_seed(43),
).images[0]
image.save("out.png")
```
完整可執行範例:[`inference_int4.py`](inference_int4.py) 批次生成 + benchmark:[`generate5.py`](generate5.py)
## 推論參數建議
| 參數 | 建議值 | 說明 |
| --- | --- | --- |
| `num_inference_steps` | `40` | 蒸餾 / 推薦步數。**請勿隨意增加**。 |
| `true_cfg_scale` | `1.0` | CFG 設定;Turbo / 蒸餾模型通常為 `0.0` 或 `1.0`(即不啟用)。 |
| `shift` | 取自 `scheduler/scheduler_config.json` | **不需手動傳入**,載入時自動套用。 |
| `height` / `width` | `1024` | 實測解析度。 |
| `compile=True` | 開啟 | 編譯模型以取得較佳效能。 |
## 範例結果
> 全部為 **40 steps / true_cfg_scale 1.0 / 1024×1024 / CPU**,seed 42–46 固定,可完全重現。
> 另附 512px 對照圖(`outputs/*_512.png`)。
> **⚠️ 512px 為縮圖**:這批 512px 圖是 1024px 輸出的 LANCZOS 縮圖,**不是**重新以 512px 推理的結果,因此沒有對應的獨立耗時資料。
### 01_hanfu — seed 42 — 980.8 s
```text
Young Chinese woman in red Hanfu, intricate embroidery, impeccable makeup, red floral forehead pattern, elaborate high bun, golden phoenix headdress, soft-lit outdoor night background, silhouetted tiered pagoda, blurred colorful distant lights, photorealistic, ultra detailed, 8k
```

### 02_astronaut — seed 43 — 865.7 s
```text
Astronaut in a jungle, cold color palette, muted colors, detailed, 8k, photorealistic, cinematic lighting
```

### 03_taipei — seed 44 — 871.5 s
```text
Cyberpunk street in Taipei at night, heavy rain, neon signs with text 'TAIPEI' and Chinese characters '台北', reflections on wet asphalt, crowded night market, cinematic, ultra detailed
```

### 04_shiba — seed 45 — 832.1 s
```text
Cute Shiba Inu wearing a tiny astronaut helmet, sitting in a field of sunflowers under a starry sky, dreamy illustration, vibrant colors, high quality
```

### 05_ink — seed 46 — 798.9 s
```text
Traditional Chinese ink wash landscape, misty mountains, a small pagoda on a cliff, cranes flying, minimalist, elegant, high aesthetic quality
```

## 效能實測摘要
完整逐 step 數據見 [`REPORT.md`](REPORT.md) 與 [`outputs/benchmark.json`](outputs/benchmark.json)。
**測試環境**
| 項目 | 內容 |
| --- | --- |
| CPU | Intel(R) Xeon(R) Platinum 8559C |
| 拓撲 | 2 sockets × 48 cores × 2 threads/core = **192 vCPU**(96 實體核心) |
| RAM | 2.0 TiB |
| 虛擬化 | KVM(完整虛擬化) |
| OpenVINO | CPU only,2026.4.0(build `2026.4.0-22959-99c81491cc3-releases/2026/4`) |
| 設定 | `num_inference_steps=40`、`true_cfg_scale=1.0`、1024×1024 |
**總結**
| 指標 | 數值 |
| --- | --- |
| 解析度 | 1024×1024 |
| 平均總耗時 | **869.80 s / 張** |
| 平均單步耗時 | 21.57 s |
| 最快 / 最慢 | 798.94 s / 980.80 s |
| 總計(5 張) | 4349.00 s |
| 記憶體高水位 | 40,211 MB |
- **512px 沒有獨立耗時資料**:`outputs/*_512.png` 是 1024px 輸出的縮圖,非重新推理。
- 文字編碼與 VAE decode 的時間**已包含**在總耗時內。
**逐張結果(1024×1024)**
| # | Prompt | Seed | 總耗時 (s) | 平均單步 (s) | true_cfg_scale |
| --- | --- | --- | --- | --- | --- |
| 01_hanfu | 42 | 980.80 | 24.257 | 1.0 |
| 02_astronaut | 43 | 865.67 | 21.497 | 1.0 |
| 03_taipei | 44 | 871.46 | 21.645 | 1.0 |
| 04_shiba | 45 | 832.12 | 20.592 | 1.0 |
| 05_ink | 46 | 798.94 | 19.846 | 1.0 |
| | **平均** | **869.80** | **21.568** | |
**模型大小**
以下為 repo 內 `openvino_model.bin` 的**實際位元組數**(Git LFS 記錄值)。
| 元件 | 位元組 | 大小 | 精度 |
| --- | ---: | ---: | --- |
| `text_encoder` | 4,256,132,893 | 4.26 GB | INT4 |
| `text_encoder_i2i` | 4,256,132,733 | 4.26 GB | INT4 |
| `transformer` | 3,696,679,634 | 3.70 GB | INT4 |
| `vision_encoder` | 577,780,620 | 0.58 GB | — |
| `vae_decoder` | 253,235,856 | 0.25 GB | — |
| `vae_encoder` | 78,002,698 | 0.08 GB | — |
| **合計** | **13,117,964,434** | **13.12 GB** | |
FP16 匯出模型約 **43.6 GB**(transformer 14.23 GB + text_encoder 15.14 GB + text_encoder_i2i 15.14 GB + vision_encoder 1.1 GB + vae_decoder 484 MB + vae_encoder 150 MB)——取自原始轉換紀錄。
INT8 部分(vision encoder + VAE)為 **0.88 GB**,佔 INT4 總量的 7%。
`text_encoder` 與 `text_encoder_i2i` 各自 INT4 後仍有 4.26 GB,是整體壓縮率(~70%)的主要限制來源。
## 檔案結構
```text
.
├── README.md # 本文件
├── REPORT.md # 完整轉換 + 實測報告
├── model_index.json # diffusers pipeline 索引(QwenImage21Pipeline)
├── openvino_config.json # OpenVINO 量化設定
├── inference_int4.py # 單張推論範例
├── generate5.py # 5 組 prompt 批次生成 + benchmark
├── quantize_int4.py # FP16 OV → INT4 OV 量化腳本
├── transformer/ # INT4 QwenImage21Transformer2DModel(32 層)
├── text_encoder/ # INT4 Qwen3VLForConditionalGeneration
├── text_encoder_i2i/ # INT4 第二 text encoder(編輯用)
├── vision_encoder/ # INT8 Qwen3VLVisionModel
├── processor/ # Qwen3VLProcessor + tokenizer
├── vae_encoder/ # INT8 VAE encoder
├── vae_decoder/ # INT8 VAE decoder
├── scheduler/ # FlowMatchEulerDiscreteScheduler 設定
├── examples/ # 5 張展示圖(與 outputs 1024 相同)
└── outputs/ # 10 張實測圖 + 3 個資料檔
├── *_1024.png # 主測組(5 張)
├── *_512.png # 對照組(5 張,為 1024 縮圖)
├── benchmark.json # 逐 step 耗時 + 記憶體 + 系統資訊
├── benchmark_quantization.json # 量化設定與耗時
└── prompts.txt # 5 組 prompt 與 seed
```
## 從零復現
```bash
# 1. 匯出 FP16 OpenVINO 模型
optimum-cli export openvino \
-m Qwen/Qwen-Image-2.1 \
--task text-to-image \
--library diffusers \
--weight-format fp16 \
./Qwen-Image-2.1-ov-fp16
# 2. NNCF weight-only INT4 量化
python quantize_int4.py --fp16-dir ./Qwen-Image-2.1-ov-fp16 \
--int4-dir ./Qwen-Image-2.1-ov-int4
# 3. 單張推論
python inference_int4.py
# 4. 批次生成 5 組 + benchmark
python generate5.py --outdir outputs
```
量化設定:
```python
from optimum.intel.openvino.configuration import (
OVConfig, OVWeightQuantizationConfig, OVPipelineQuantizationConfig,
)
int4_config = OVWeightQuantizationConfig(
bits=4, sym=False, group_size=128,
group_size_fallback="adjust", ratio=1.0,
)
pipeline_config = OVPipelineQuantizationConfig(
quantization_configs={
"transformer": int4_config,
"text_encoder": int4_config,
"text_encoder_i2i": int4_config,
},
default_config=OVWeightQuantizationConfig(bits=8),
)
```
## 已知限制
- **CPU 速度最慢**:1024×1024 / 40 steps 平均 **870 s / 張**(約 14.5 分鐘),是這批 repo 中最慢的,純 CPU 實務上不太實用。
- **記憶體需求最高**:INT4 模型 13.12 GB,實測行程結束時 RSS 達 **40.2 GB**,建議至少預留 48 GB 可用記憶體。
- **需要 diffusers 主分支**:`diffusers==0.41.0.dev0`(含 `QwenImage21Pipeline`,v0.40.0 之後才加入)。`optimum-intel` 需使用 `2.3.0.dev0+32a317a`。
- **需要對上游套件套用兩處修改**:
- 1. `transformers` 的相依套件上限(hub cap `<1.0` → `<2.0`)
- 2. `optimum-intel` 的 `modeling_visual_language.py`(transformers ≥ 5 的 `VisionRotaryEmbedding` 別名)
- **512px 對照為縮圖**:`outputs/*_512.png` 是 1024px 輸出的 LANCZOS 縮圖,非重新推理。
- **`model_index.json` 未列出 `text_encoder_i2i` 與 `vision_encoder`**:這兩個目錄存在於 repo 中但未在索引中宣告,屬 optimum-intel 匯出的已知落差。
- **10 張風格展示圖無實測紀錄**:只有 5 張主測圖有 benchmark 資料。
## 授權與出處
- **來源模型**:[`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1)
- **授權**:`qwen-research`(`license: other`)
- **完整條款**:[LICENSE](https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE)
- **轉換**:僅做格式轉換與權重量化,模型權重來自來源模型
使用本模型時請遵守來源模型的授權條款。
---
<div align="center">
**Made with OpenVINO + optimum-intel + NNCF**
</div> |