# Qwen-Image-2.1 — OpenVINO INT4:轉換與實測報告 | 項目 | 內容 | | --- | --- | | **來源模型** | [`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1) | | **Pipeline 類型** | `QwenImage21Pipeline` | | **Scheduler** | `FlowMatchEulerDiscreteScheduler`(`shift=1.0`、`shift_terminal=0.02`、`use_dynamic_shifting=true`) | | **目標格式** | OpenVINO IR;Transformer 與兩個 Qwen3-VL text encoder 為 INT4,vision encoder 與 VAE 為 INT8 | | **發布位置** | [`HelloSun/Qwen-Image-2.1-OpenVINO-INT4`](https://huggingface.co/HelloSun/Qwen-Image-2.1-OpenVINO-INT4) | --- ## 目錄 1. 環境與版本 2. 轉換流程 3. 模型檔案與體積 4. 實測結果 5. 測試 Prompt 6. 觀察與結論 7. 復現步驟 ## 1. 環境與版本 ```text diffusers==0.41.0.dev0 transformers==5.10.4 tokenizers==0.22.2 huggingface-hub==1.33.0 optimum==2.3.0 optimum-intel==2.3.0.dev0+32a317a openvino==2026.4.0 nncf==3.4.0 torch==2.14.0 pillow==12.3.0 psutil==7.2.2 ``` ## 2. 轉換流程 ### 步驟 1 — 匯出 OpenVINO FP16 ```bash optimum-cli export openvino -m Qwen/Qwen-Image-2.1 --task text-to-image --library diffusers --weight-format fp16 ./Qwen-Image-2.1-ov-fp16 ``` ### 步驟 2 — NNCF weight-only 量化 由 [`quantize_int4.py`](quantize_int4.py) 執行。 ```python from optimum.intel.openvino.configuration import ( OVConfig, OVWeightQuantizationConfig, OVPipelineQuantizationConfig, ) int4_config = OVWeightQuantizationConfig( bits=4, sym=False, group_size=128, group_size_fallback="adjust", ratio=1.0, ) pipeline_config = OVPipelineQuantizationConfig( quantization_configs={ "transformer": int4_config, "text_encoder": int4_config, "text_encoder_i2i": int4_config, }, default_config=OVWeightQuantizationConfig(bits=8), ) ``` **重點說明** - `bits=4`、`sym=False` → **非對稱 4-bit** weight-only 量化。 - `group_size=128` → 每 128 個權重共用一組 scale / zero-point。 - `group_size_fallback="adjust"` → 分組不整除時自動調整群組大小,避免部分層被跳過。 - `ratio=1.0` → 全部權重都量化,不做混合精度保留。 - 未被列出的元件(VAE、vision encoder 等)沿用 `default_config` 的 INT8。 - **注意**:`openvino_config.json` 的 `default_config` 只記錄 `quant_method`,未寫入 `bits=8`;INT8 fallback 只能從 `quantize_int4.py` 得知。 ## 3. 模型檔案與體積 | 元件 | 精度 | 說明 | | --- | --- | --- | | `transformer` | **INT4** | | | `text_encoder` | **INT4** | | | `text_encoder_i2i` | **INT4** | | | `vae_decoder` | INT8 | 由 latent 解回 RGB | | `vae_encoder` | INT8 | image encoder(推論時未實際使用) | | `tokenizer / processor` | — | 未量化 | | `scheduler` | — | 設定取自 `scheduler/scheduler_config.json` | ### 實際體積 下表為 repo 內 `openvino_model.bin` 的實際位元組數(Git LFS 記錄值),可直接對照下載量。 | 元件 | 位元組 | GB | GiB | | --- | ---: | ---: | ---: | | `text_encoder` | 4,256,132,893 | 4.26 GB | 3.96 GiB | | `text_encoder_i2i` | 4,256,132,733 | 4.26 GB | 3.96 GiB | | `transformer` | 3,696,679,634 | 3.70 GB | 3.44 GiB | | `vision_encoder` | 577,780,620 | 0.58 GB | 0.54 GiB | | `vae_decoder` | 253,235,856 | 0.25 GB | 0.24 GiB | | `vae_encoder` | 78,002,698 | 0.08 GB | 0.07 GiB | | **合計** | **13,117,964,434** | **13.12 GB** | **12.22 GiB** | FP16 匯出模型約 **43.6 GB**(transformer 14.23 GB + text_encoder 15.14 GB + text_encoder_i2i 15.14 GB + vision_encoder 1.1 GB + vae_decoder 484 MB + vae_encoder 150 MB)——取自原始轉換紀錄。 INT8 部分(vision encoder + VAE)為 **0.88 GB**,佔 INT4 總量的 7%。 `text_encoder` 與 `text_encoder_i2i` 各自 INT4 後仍有 4.26 GB,是整體壓縮率(~70%)的主要限制來源。 ## 4. 實測結果 | 項目 | 數值 | | --- | --- | | 樣本數 | 5 | | 推論設定 | `num_inference_steps=40`、`true_cfg_scale=1.0`、1024×1024 | | `from_pretrained` + `compile=True` | **14.69 s** | | 平均總耗時 | **869.80 s** | | 平均單步耗時 | 21.57 s | | 記憶體高水位 | 40,210.9 MB | ### 逐張結果 | # | Prompt | Seed | steps | CFG | 總耗時 (s) | 平均單步 (s) | 逐 step 耗時 (s) | | --- | --- | --- | --- | --- | --- | --- | --- | | 01_hanfu | 42 | 40 | true_cfg_scale=1.0 | 980.80 | 24.257 | 25.11, 25.28, 26.13, 24.27, 23.42, 23.69, 23.61, 24.40, 24.30, 23.40, 23.70, 24.30, 25.10, 24.08, 23.92, 23.29, 22.89, 22.52, 23.98, 23.21, 22.70, 22.82, 23.98, 26.02, 25.48, 24.72, 25.46, 26.10, 26.74, 24.78, 24.33, 23.99, 23.37, 23.93, 24.49, 23.70, 23.89, 24.01, 24.92 | | 02_astronaut | 43 | 40 | true_cfg_scale=1.0 | 865.67 | 21.497 | 21.50, 21.01, 21.73, 21.86, 22.20, 21.60, 21.19, 22.10, 21.79, 21.39, 21.83, 21.90, 20.80, 21.49, 21.40, 21.80, 20.89, 21.30, 21.30, 21.81, 21.77, 22.56, 22.54, 22.00, 23.34, 20.47, 20.81, 20.30, 20.19, 20.51, 20.81, 21.29, 20.70, 21.29, 22.22, 21.48, 21.61, 21.68, 21.92 | | 03_taipei | 44 | 40 | true_cfg_scale=1.0 | 871.46 | 21.645 | 20.98, 21.01, 21.49, 21.20, 21.76, 22.05, 21.59, 20.72, 21.98, 21.10, 22.48, 21.11, 21.33, 21.68, 22.21, 21.67, 19.91, 20.10, 20.30, 21.11, 21.49, 20.00, 21.61, 21.78, 20.51, 21.11, 20.80, 20.49, 20.60, 20.00, 19.39, 20.21, 24.41, 25.51, 26.02, 25.39, 22.58, 21.80, 24.70 | | 04_shiba | 45 | 40 | true_cfg_scale=1.0 | 832.12 | 20.592 | 23.89, 22.51, 20.41, 20.30, 20.21, 20.00, 20.28, 19.71, 19.60, 19.83, 19.79, 20.39, 20.00, 20.22, 20.08, 19.80, 20.01, 19.89, 19.79, 20.10, 19.80, 20.02, 19.67, 19.40, 22.24, 24.57, 25.12, 25.00, 20.57, 20.10, 20.82, 21.00, 19.80, 20.20, 19.92, 19.30, 19.46, 19.80, 19.49 | | 05_ink | 46 | 40 | true_cfg_scale=1.0 | 798.94 | 19.846 | 20.20, 19.91, 20.10, 19.60, 20.79, 19.80, 19.90, 19.41, 19.29, 19.70, 19.71, 19.69, 19.49, 19.40, 19.51, 20.29, 19.49, 19.40, 19.81, 20.50, 20.03, 19.87, 19.91, 19.39, 19.91, 19.90, 20.01, 20.29, 20.19, 20.00, 19.69, 20.21, 19.60, 20.51, 19.98, 19.50, 19.60, 19.69, 19.71 | | **平均** | | | | **869.80** | **21.568** | | ### 逐 step 耗時(主測組) | # | Prompt | Seed | 總耗時 (s) | 平均單步 (s) | 逐 step 耗時 (s) | | --- | --- | --- | ---: | ---: | --- | | 01 | hanfu | 42 | 980.80 | 24.257 | 25.11, 25.28, 26.13, 24.27, 23.42, 23.69, 23.61, 24.40, 24.30, 23.40, 23.70, 24.30, 25.10, 24.08, 23.92, 23.29, 22.89, 22.52, 23.98, 23.21, 22.70, 22.82, 23.98, 26.02, 25.48, 24.72, 25.46, 26.10, 26.74, 24.78, 24.33, 23.99, 23.37, 23.93, 24.49, 23.70, 23.89, 24.01, 24.92 | | 02 | astronaut | 43 | 865.67 | 21.497 | 21.50, 21.01, 21.73, 21.86, 22.20, 21.60, 21.19, 22.10, 21.79, 21.39, 21.83, 21.90, 20.80, 21.49, 21.40, 21.80, 20.89, 21.30, 21.30, 21.81, 21.77, 22.56, 22.54, 22.00, 23.34, 20.47, 20.81, 20.30, 20.19, 20.51, 20.81, 21.29, 20.70, 21.29, 22.22, 21.48, 21.61, 21.68, 21.92 | | 03 | taipei | 44 | 871.46 | 21.645 | 20.98, 21.01, 21.49, 21.20, 21.76, 22.05, 21.59, 20.72, 21.98, 21.10, 22.48, 21.11, 21.33, 21.68, 22.21, 21.67, 19.91, 20.10, 20.30, 21.11, 21.49, 20.00, 21.61, 21.78, 20.51, 21.11, 20.80, 20.49, 20.60, 20.00, 19.39, 20.21, 24.41, 25.51, 26.02, 25.39, 22.58, 21.80, 24.70 | | 04 | shiba | 45 | 832.12 | 20.592 | 23.89, 22.51, 20.41, 20.30, 20.21, 20.00, 20.28, 19.71, 19.60, 19.83, 19.79, 20.39, 20.00, 20.22, 20.08, 19.80, 20.01, 19.89, 19.79, 20.10, 19.80, 20.02, 19.67, 19.40, 22.24, 24.57, 25.12, 25.00, 20.57, 20.10, 20.82, 21.00, 19.80, 20.20, 19.92, 19.30, 19.46, 19.80, 19.49 | | 05 | ink | 46 | 798.94 | 19.846 | 20.20, 19.91, 20.10, 19.60, 20.79, 19.80, 19.90, 19.41, 19.29, 19.70, 19.71, 19.69, 19.49, 19.40, 19.51, 20.29, 19.49, 19.40, 19.81, 20.50, 20.03, 19.87, 19.91, 19.39, 19.91, 19.90, 20.01, 20.29, 20.19, 20.00, 19.69, 20.21, 19.60, 20.51, 19.98, 19.50, 19.60, 19.69, 19.71 | > 「總耗時」包含文字編碼與 VAE decode;「平均單步」= 逐 step 耗時加總 ÷ 步數, > 因此會略小於 `總耗時 ÷ 步數`。 ## 5. 測試 Prompt | # | Seed | Prompt | | --- | --- | --- | | 01 | 42 | `Young Chinese woman in red Hanfu, intricate embroidery, impeccable makeup, red floral forehead pattern, elaborate high bun, golden phoenix headdress, soft-lit outdoor night background, silhouetted tiered pagoda, blurred colorful distant lights, photorealistic, ultra detailed, 8k` | | 02 | 43 | `Astronaut in a jungle, cold color palette, muted colors, detailed, 8k, photorealistic, cinematic lighting` | | 03 | 44 | `Cyberpunk street in Taipei at night, heavy rain, neon signs with text 'TAIPEI' and Chinese characters '台北', reflections on wet asphalt, crowded night market, cinematic, ultra detailed` | | 04 | 45 | `Cute Shiba Inu wearing a tiny astronaut helmet, sitting in a field of sunflowers under a starry sky, dreamy illustration, vibrant colors, high quality` | | 05 | 46 | `Traditional Chinese ink wash landscape, misty mountains, a small pagoda on a cliff, cranes flying, minimalist, elegant, high aesthetic quality` | 完整內容另見 [`outputs/prompts.txt`](outputs/prompts.txt)。 ## 6. 觀察與結論 1. **速度瓶颈在步數**:40 steps × 平均 21.6 s/步,是這批 repo 中單步最慢的(transformer 4B 級別但架構較複雜,且帶 Qwen3-VL 語言條件)。 2. **第一張的額外開銷最明顯**:01_hanfu 為 980.80 s,之後四張降至 798.94–871.46 s,首張比後續平均高出約 16%,warmup 效應顯著。 3. **記憶體在第二張後達平台**:peak RSS 從 36.2 GB 升至 40.0 GB 後維持不變。 4. **`step_times` 只有 39 個值而非 40 個**:計時以第一次 callback 為基準,第一個 step 的時間被當作起始基準扣除。 5. **INT8 覆蓋率低**:只有 vision encoder 與 VAE 共 0.88 GB 走 INT8,97% 的權重都經過 INT4 量化,因此整體壓縮率(3.3×)低於 FLUX.2 的 3.6×。 ## 7. 復現步驟 ```bash # 1. 匯出 FP16 OpenVINO 模型 optimum-cli export openvino \ -m Qwen/Qwen-Image-2.1 \ --task text-to-image \ --library diffusers \ --weight-format fp16 \ ./Qwen-Image-2.1-ov-fp16 # 2. NNCF weight-only INT4 量化 python quantize_int4.py --fp16-dir ./Qwen-Image-2.1-ov-fp16 \ --int4-dir ./Qwen-Image-2.1-ov-int4 # 3. 單張推論 python inference_int4.py # 4. 批次生成 5 組 + benchmark python generate5.py --outdir outputs ``` 量化腳本會另外輸出: - [`outputs/benchmark.json`](outputs/benchmark.json) — 逐 step 耗時、記憶體、系統資訊 - [`outputs/prompts.txt`](outputs/prompts.txt) — prompt 與 seed 清單 - `outputs/*.png` — 主測組 1024px 與對照組 512px --- 原始機器紀錄為 [`outputs/benchmark.json`](outputs/benchmark.json)。