File size: 12,343 Bytes
4825e8b
 
 
 
cbe716d
 
4825e8b
 
 
 
 
 
 
 
 
cbe716d
4825e8b
cbe716d
4825e8b
cbe716d
4825e8b
cbe716d
 
 
 
 
4825e8b
cbe716d
4825e8b
cbe716d
4825e8b
cbe716d
 
 
 
 
 
 
 
 
 
 
4825e8b
cbe716d
4825e8b
cbe716d
4825e8b
cbe716d
 
 
 
 
 
 
 
 
 
4825e8b
cbe716d
4825e8b
cbe716d
 
 
 
 
 
4825e8b
cbe716d
4825e8b
 
cbe716d
4825e8b
 
cbe716d
 
 
 
4825e8b
 
cbe716d
 
4825e8b
cbe716d
 
 
 
 
 
 
 
 
 
 
4825e8b
 
cbe716d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4825e8b
 
cbe716d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4825e8b
 
cbe716d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4825e8b
cbe716d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4825e8b
cbe716d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
---
license: other
license_name: qwen-research
license_link: https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE
library_name: optimum-intel
pipeline_tag: text-to-image
base_model: Qwen/Qwen-Image-2.1
tags:
- openvino
- optimum-intel
- int4
- nncf
- text-to-image
---

# Qwen-Image-2.1 — OpenVINO INT4

將 [`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1) 轉換為 **OpenVINO INT4** 權重,並以 **純 CPU** 實測出圖。

`optimum-intel` 匯出 → NNCF weight-only INT4 量化(耗時 115 s)→ **FP16 43.6 GB → INT4 13.12 GB(↓ ~70%,約 3.3×)**。

> ### ⚠️ 環境需求與其他 repo 不同
>
> `QwenImage21Pipeline` 在 diffusers **0.40.0 之後**才加入,因此 `diffusers==0.37.1` **無法使用**。
> 本 repo 需要 diffusers 主分支版本,且需對 `transformers` 與 `optimum-intel` 套用兩處小幅度修改
>(詳見下方安裝說明與 `REPORT.md`)。

---

## 目錄

- [快速資訊](#快速資訊)
- [特色](#特色)
- [安裝](#安裝)
- [快速開始](#快速開始)
- [推論參數建議](#推論參數建議)
- [範例結果](#範例結果)
- [效能實測摘要](#效能實測摘要)
- [檔案結構](#檔案結構)
- [從零復現](#從零復現)
- [已知限制](#已知限制)
- [授權與出處](#授權與出處)

---

## 快速資訊

| 項目 | 內容 |
| --- | --- |
| **基座模型** | [`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1) |
| **Pipeline** | `QwenImage21Pipeline` |
| **Scheduler** | `FlowMatchEulerDiscreteScheduler`(`shift=1.0`、`shift_terminal=0.02`、`use_dynamic_shifting=true`) |
| **量化格式** | `transformer`、`text_encoder`、`text_encoder_i2i` → INT4;其餘 → INT8 |
| **模型大小** | FP16 43.6 GB → **INT4 13.12 GB**(↓ ~70%,約 3.3×) |
| **推論裝置** | CPU(OpenVINO CPU plugin,無需 GPU) |
| **推薦參數** | `num_inference_steps=40`、`true_cfg_scale=1.0`、`1024×1024` |
| **轉換工具** | optimum-intel 2.2.x / optimum 2.3.0 / OpenVINO 2026.4.0 / NNCF 3.4.0 |

## 特色

- **三個元件 INT4**:Transformer 與兩個 Qwen3-VL text encoder 全部 INT4,`text_encoder_i2i` 是編輯功能用的第二個 text encoder。
- **⚠️ 需要 diffusers 主分支**:`QwenImage21Pipeline` 需要 `diffusers==0.41.0.dev0`與 `transformers==5.10.4`,與同批其他 repo 的版本需求不同。
- **體積 ↓ 70%**:FP16 43.6 GB → INT4 13.12 GB(約 3.3×)。
- **支援編輯能力**:保留了 image-to-image 的 `text_encoder_i2i` 與 `vision_encoder`。
- **⚠️ 速度最慢**:1024×1024 / 40 steps 平均 870 s / 張(約 14.5 分鐘)。
- **記憶體需求最高**:行程結束時 RSS 達 40.2 GB。

## 安裝

```bash
pip install diffusers==0.41.0.dev0 transformers==5.10.4 tokenizers==0.22.2 huggingface-hub==1.33.0 optimum==2.3.0 optimum-intel==2.3.0.dev0+32a317a openvino==2026.4.0 nncf==3.4.0 torch==2.14.0 pillow==12.3.0 psutil==7.2.2
```

`nncf` 僅重新量化時需要。

## 快速開始

```python
import torch
from optimum.intel import OVDiffusionPipeline

pipe = OVDiffusionPipeline.from_pretrained("HelloSun/Qwen-Image-2.1-OpenVINO-INT4", compile=True)

image = pipe(
    prompt="Astronaut in a jungle, cold color palette, muted colors, "
             "detailed, 8k, photorealistic, cinematic lighting",
    height=1024, width=1024,
    num_inference_steps=40,
    true_cfg_scale=1.0,
    generator=torch.Generator().manual_seed(43),
).images[0]

image.save("out.png")
```

完整可執行範例:[`inference_int4.py`](inference_int4.py) 批次生成 + benchmark:[`generate5.py`](generate5.py)

## 推論參數建議

| 參數 | 建議值 | 說明 |
| --- | --- | --- |
| `num_inference_steps` | `40` | 蒸餾 / 推薦步數。**請勿隨意增加**。 |
| `true_cfg_scale` | `1.0` | CFG 設定;Turbo / 蒸餾模型通常為 `0.0` 或 `1.0`(即不啟用)。 |
| `shift` | 取自 `scheduler/scheduler_config.json` | **不需手動傳入**,載入時自動套用。 |
| `height` / `width` | `1024` | 實測解析度。 |
| `compile=True` | 開啟 | 編譯模型以取得較佳效能。 |

## 範例結果

> 全部為 **40 steps / true_cfg_scale 1.0 / 1024×1024 / CPU**,seed 42–46 固定,可完全重現。

> 另附 512px 對照圖(`outputs/*_512.png`)。

> **⚠️ 512px 為縮圖**:這批 512px 圖是 1024px 輸出的 LANCZOS 縮圖,**不是**重新以 512px 推理的結果,因此沒有對應的獨立耗時資料。

### 01_hanfu — seed 42 — 980.8 s

```text
Young Chinese woman in red Hanfu, intricate embroidery, impeccable makeup, red floral forehead pattern, elaborate high bun, golden phoenix headdress, soft-lit outdoor night background, silhouetted tiered pagoda, blurred colorful distant lights, photorealistic, ultra detailed, 8k
```

![01_hanfu](outputs/01_hanfu_1024.png)

### 02_astronaut — seed 43 — 865.7 s

```text
Astronaut in a jungle, cold color palette, muted colors, detailed, 8k, photorealistic, cinematic lighting
```

![02_astronaut](outputs/02_astronaut_1024.png)

### 03_taipei — seed 44 — 871.5 s

```text
Cyberpunk street in Taipei at night, heavy rain, neon signs with text 'TAIPEI' and Chinese characters '台北', reflections on wet asphalt, crowded night market, cinematic, ultra detailed
```

![03_taipei](outputs/03_taipei_1024.png)

### 04_shiba — seed 45 — 832.1 s

```text
Cute Shiba Inu wearing a tiny astronaut helmet, sitting in a field of sunflowers under a starry sky, dreamy illustration, vibrant colors, high quality
```

![04_shiba](outputs/04_shiba_1024.png)

### 05_ink — seed 46 — 798.9 s

```text
Traditional Chinese ink wash landscape, misty mountains, a small pagoda on a cliff, cranes flying, minimalist, elegant, high aesthetic quality
```

![05_ink](outputs/05_ink_1024.png)

## 效能實測摘要

完整逐 step 數據見 [`REPORT.md`](REPORT.md) 與 [`outputs/benchmark.json`](outputs/benchmark.json)。

**測試環境**

| 項目 | 內容 |
| --- | --- |
| CPU | Intel(R) Xeon(R) Platinum 8559C |
| 拓撲 | 2 sockets × 48 cores × 2 threads/core = **192 vCPU**(96 實體核心) |
| RAM | 2.0 TiB |
| 虛擬化 | KVM(完整虛擬化) |
| OpenVINO | CPU only,2026.4.0(build `2026.4.0-22959-99c81491cc3-releases/2026/4`) |
| 設定 | `num_inference_steps=40`、`true_cfg_scale=1.0`、1024×1024 |

**總結**

| 指標 | 數值 |
| --- | --- |
| 解析度 | 1024×1024 |
| 平均總耗時 | **869.80 s / 張** |
| 平均單步耗時 | 21.57 s |
| 最快 / 最慢 | 798.94 s / 980.80 s |
| 總計(5 張) | 4349.00 s |
| 記憶體高水位 | 40,211 MB |

- **512px 沒有獨立耗時資料**:`outputs/*_512.png` 是 1024px 輸出的縮圖,非重新推理。
- 文字編碼與 VAE decode 的時間**已包含**在總耗時內。

**逐張結果(1024×1024)**

| # | Prompt | Seed | 總耗時 (s) | 平均單步 (s) | true_cfg_scale |
| --- | --- | --- | --- | --- | --- |
| 01_hanfu | 42 | 980.80 | 24.257 | 1.0 |
| 02_astronaut | 43 | 865.67 | 21.497 | 1.0 |
| 03_taipei | 44 | 871.46 | 21.645 | 1.0 |
| 04_shiba | 45 | 832.12 | 20.592 | 1.0 |
| 05_ink | 46 | 798.94 | 19.846 | 1.0 |
|  | **平均** | **869.80** | **21.568** |  |

**模型大小**

以下為 repo 內 `openvino_model.bin` 的**實際位元組數**(Git LFS 記錄值)。

| 元件 | 位元組 | 大小 | 精度 |
| --- | ---: | ---: | --- |
| `text_encoder` | 4,256,132,893 | 4.26 GB | INT4 |
| `text_encoder_i2i` | 4,256,132,733 | 4.26 GB | INT4 |
| `transformer` | 3,696,679,634 | 3.70 GB | INT4 |
| `vision_encoder` | 577,780,620 | 0.58 GB | — |
| `vae_decoder` | 253,235,856 | 0.25 GB | — |
| `vae_encoder` | 78,002,698 | 0.08 GB | — |
| **合計** | **13,117,964,434** | **13.12 GB** | |

FP16 匯出模型約 **43.6 GB**(transformer 14.23 GB + text_encoder 15.14 GB + text_encoder_i2i 15.14 GB + vision_encoder 1.1 GB + vae_decoder 484 MB + vae_encoder 150 MB)——取自原始轉換紀錄。

INT8 部分(vision encoder + VAE)為 **0.88 GB**,佔 INT4 總量的 7%。

`text_encoder` 與 `text_encoder_i2i` 各自 INT4 後仍有 4.26 GB,是整體壓縮率(~70%)的主要限制來源。

## 檔案結構

```text
.
├── README.md                  # 本文件
├── REPORT.md                  # 完整轉換 + 實測報告
├── model_index.json           # diffusers pipeline 索引(QwenImage21Pipeline)
├── openvino_config.json       # OpenVINO 量化設定
├── inference_int4.py          # 單張推論範例
├── generate5.py               # 5 組 prompt 批次生成 + benchmark
├── quantize_int4.py           # FP16 OV → INT4 OV 量化腳本
├── transformer/               # INT4 QwenImage21Transformer2DModel(32 層)
├── text_encoder/              # INT4 Qwen3VLForConditionalGeneration
├── text_encoder_i2i/          # INT4 第二 text encoder(編輯用)
├── vision_encoder/            # INT8 Qwen3VLVisionModel
├── processor/                 # Qwen3VLProcessor + tokenizer
├── vae_encoder/               # INT8 VAE encoder
├── vae_decoder/               # INT8 VAE decoder
├── scheduler/                 # FlowMatchEulerDiscreteScheduler 設定
├── examples/                  # 5 張展示圖(與 outputs 1024 相同)
└── outputs/                   # 10 張實測圖 + 3 個資料檔
    ├── *_1024.png             # 主測組(5 張)
    ├── *_512.png              # 對照組(5 張,為 1024 縮圖)
    ├── benchmark.json         # 逐 step 耗時 + 記憶體 + 系統資訊
    ├── benchmark_quantization.json  # 量化設定與耗時
    └── prompts.txt            # 5 組 prompt 與 seed
```

## 從零復現

```bash
# 1. 匯出 FP16 OpenVINO 模型
optimum-cli export openvino \
  -m Qwen/Qwen-Image-2.1 \
  --task text-to-image \
  --library diffusers \
  --weight-format fp16 \
  ./Qwen-Image-2.1-ov-fp16

# 2. NNCF weight-only INT4 量化
python quantize_int4.py --fp16-dir ./Qwen-Image-2.1-ov-fp16 \
                          --int4-dir  ./Qwen-Image-2.1-ov-int4

# 3. 單張推論
python inference_int4.py

# 4. 批次生成 5 組 + benchmark
python generate5.py --outdir outputs
```

量化設定:

```python
from optimum.intel.openvino.configuration import (
    OVConfig, OVWeightQuantizationConfig, OVPipelineQuantizationConfig,
)

int4_config = OVWeightQuantizationConfig(
    bits=4, sym=False, group_size=128,
    group_size_fallback="adjust", ratio=1.0,
)

pipeline_config = OVPipelineQuantizationConfig(
    quantization_configs={
        "transformer":  int4_config,
        "text_encoder": int4_config,
        "text_encoder_i2i": int4_config,
    },
    default_config=OVWeightQuantizationConfig(bits=8),
)
```

## 已知限制

- **CPU 速度最慢**:1024×1024 / 40 steps 平均 **870 s / 張**(約 14.5 分鐘),是這批 repo 中最慢的,純 CPU 實務上不太實用。
- **記憶體需求最高**:INT4 模型 13.12 GB,實測行程結束時 RSS 達 **40.2 GB**,建議至少預留 48 GB 可用記憶體。
- **需要 diffusers 主分支**:`diffusers==0.41.0.dev0`(含 `QwenImage21Pipeline`,v0.40.0 之後才加入)。`optimum-intel` 需使用 `2.3.0.dev0+32a317a`。
- **需要對上游套件套用兩處修改**:
-   1. `transformers` 的相依套件上限(hub cap `<1.0` → `<2.0`)
-   2. `optimum-intel` 的 `modeling_visual_language.py`(transformers ≥ 5 的 `VisionRotaryEmbedding` 別名)

- **512px 對照為縮圖**:`outputs/*_512.png` 是 1024px 輸出的 LANCZOS 縮圖,非重新推理。
- **`model_index.json` 未列出 `text_encoder_i2i` 與 `vision_encoder`**:這兩個目錄存在於 repo 中但未在索引中宣告,屬 optimum-intel 匯出的已知落差。
- **10 張風格展示圖無實測紀錄**:只有 5 張主測圖有 benchmark 資料。

## 授權與出處

- **來源模型**:[`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1)

- **授權**:`qwen-research`(`license: other`)

- **完整條款**:[LICENSE](https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE)

- **轉換**:僅做格式轉換與權重量化,模型權重來自來源模型

使用本模型時請遵守來源模型的授權條款。

---

<div align="center">

**Made with OpenVINO + optimum-intel + NNCF**

</div>