--- license: other license_name: qwen-research license_link: https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE library_name: optimum-intel pipeline_tag: text-to-image base_model: Qwen/Qwen-Image-2.1 tags: - openvino - optimum-intel - int4 - nncf - text-to-image --- # Qwen-Image-2.1 — OpenVINO INT4 將 [`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1) 轉換為 **OpenVINO INT4** 權重,並以 **純 CPU** 實測出圖。 `optimum-intel` 匯出 → NNCF weight-only INT4 量化(耗時 115 s)→ **FP16 43.6 GB → INT4 13.12 GB(↓ ~70%,約 3.3×)**。 > ### ⚠️ 環境需求與其他 repo 不同 > > `QwenImage21Pipeline` 在 diffusers **0.40.0 之後**才加入,因此 `diffusers==0.37.1` **無法使用**。 > 本 repo 需要 diffusers 主分支版本,且需對 `transformers` 與 `optimum-intel` 套用兩處小幅度修改 >(詳見下方安裝說明與 `REPORT.md`)。 --- ## 目錄 - [快速資訊](#快速資訊) - [特色](#特色) - [安裝](#安裝) - [快速開始](#快速開始) - [推論參數建議](#推論參數建議) - [範例結果](#範例結果) - [效能實測摘要](#效能實測摘要) - [檔案結構](#檔案結構) - [從零復現](#從零復現) - [已知限制](#已知限制) - [授權與出處](#授權與出處) --- ## 快速資訊 | 項目 | 內容 | | --- | --- | | **基座模型** | [`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1) | | **Pipeline** | `QwenImage21Pipeline` | | **Scheduler** | `FlowMatchEulerDiscreteScheduler`(`shift=1.0`、`shift_terminal=0.02`、`use_dynamic_shifting=true`) | | **量化格式** | `transformer`、`text_encoder`、`text_encoder_i2i` → INT4;其餘 → INT8 | | **模型大小** | FP16 43.6 GB → **INT4 13.12 GB**(↓ ~70%,約 3.3×) | | **推論裝置** | CPU(OpenVINO CPU plugin,無需 GPU) | | **推薦參數** | `num_inference_steps=40`、`true_cfg_scale=1.0`、`1024×1024` | | **轉換工具** | optimum-intel 2.2.x / optimum 2.3.0 / OpenVINO 2026.4.0 / NNCF 3.4.0 | ## 特色 - **三個元件 INT4**:Transformer 與兩個 Qwen3-VL text encoder 全部 INT4,`text_encoder_i2i` 是編輯功能用的第二個 text encoder。 - **⚠️ 需要 diffusers 主分支**:`QwenImage21Pipeline` 需要 `diffusers==0.41.0.dev0`與 `transformers==5.10.4`,與同批其他 repo 的版本需求不同。 - **體積 ↓ 70%**:FP16 43.6 GB → INT4 13.12 GB(約 3.3×)。 - **支援編輯能力**:保留了 image-to-image 的 `text_encoder_i2i` 與 `vision_encoder`。 - **⚠️ 速度最慢**:1024×1024 / 40 steps 平均 870 s / 張(約 14.5 分鐘)。 - **記憶體需求最高**:行程結束時 RSS 達 40.2 GB。 ## 安裝 ```bash pip install diffusers==0.41.0.dev0 transformers==5.10.4 tokenizers==0.22.2 huggingface-hub==1.33.0 optimum==2.3.0 optimum-intel==2.3.0.dev0+32a317a openvino==2026.4.0 nncf==3.4.0 torch==2.14.0 pillow==12.3.0 psutil==7.2.2 ``` `nncf` 僅重新量化時需要。 ## 快速開始 ```python import torch from optimum.intel import OVDiffusionPipeline pipe = OVDiffusionPipeline.from_pretrained("HelloSun/Qwen-Image-2.1-OpenVINO-INT4", compile=True) image = pipe( prompt="Astronaut in a jungle, cold color palette, muted colors, " "detailed, 8k, photorealistic, cinematic lighting", height=1024, width=1024, num_inference_steps=40, true_cfg_scale=1.0, generator=torch.Generator().manual_seed(43), ).images[0] image.save("out.png") ``` 完整可執行範例:[`inference_int4.py`](inference_int4.py) 批次生成 + benchmark:[`generate5.py`](generate5.py) ## 推論參數建議 | 參數 | 建議值 | 說明 | | --- | --- | --- | | `num_inference_steps` | `40` | 蒸餾 / 推薦步數。**請勿隨意增加**。 | | `true_cfg_scale` | `1.0` | CFG 設定;Turbo / 蒸餾模型通常為 `0.0` 或 `1.0`(即不啟用)。 | | `shift` | 取自 `scheduler/scheduler_config.json` | **不需手動傳入**,載入時自動套用。 | | `height` / `width` | `1024` | 實測解析度。 | | `compile=True` | 開啟 | 編譯模型以取得較佳效能。 | ## 範例結果 > 全部為 **40 steps / true_cfg_scale 1.0 / 1024×1024 / CPU**,seed 42–46 固定,可完全重現。 > 另附 512px 對照圖(`outputs/*_512.png`)。 > **⚠️ 512px 為縮圖**:這批 512px 圖是 1024px 輸出的 LANCZOS 縮圖,**不是**重新以 512px 推理的結果,因此沒有對應的獨立耗時資料。 ### 01_hanfu — seed 42 — 980.8 s ```text Young Chinese woman in red Hanfu, intricate embroidery, impeccable makeup, red floral forehead pattern, elaborate high bun, golden phoenix headdress, soft-lit outdoor night background, silhouetted tiered pagoda, blurred colorful distant lights, photorealistic, ultra detailed, 8k ``` ![01_hanfu](outputs/01_hanfu_1024.png) ### 02_astronaut — seed 43 — 865.7 s ```text Astronaut in a jungle, cold color palette, muted colors, detailed, 8k, photorealistic, cinematic lighting ``` ![02_astronaut](outputs/02_astronaut_1024.png) ### 03_taipei — seed 44 — 871.5 s ```text Cyberpunk street in Taipei at night, heavy rain, neon signs with text 'TAIPEI' and Chinese characters '台北', reflections on wet asphalt, crowded night market, cinematic, ultra detailed ``` ![03_taipei](outputs/03_taipei_1024.png) ### 04_shiba — seed 45 — 832.1 s ```text Cute Shiba Inu wearing a tiny astronaut helmet, sitting in a field of sunflowers under a starry sky, dreamy illustration, vibrant colors, high quality ``` ![04_shiba](outputs/04_shiba_1024.png) ### 05_ink — seed 46 — 798.9 s ```text Traditional Chinese ink wash landscape, misty mountains, a small pagoda on a cliff, cranes flying, minimalist, elegant, high aesthetic quality ``` ![05_ink](outputs/05_ink_1024.png) ## 效能實測摘要 完整逐 step 數據見 [`REPORT.md`](REPORT.md) 與 [`outputs/benchmark.json`](outputs/benchmark.json)。 **測試環境** | 項目 | 內容 | | --- | --- | | CPU | Intel(R) Xeon(R) Platinum 8559C | | 拓撲 | 2 sockets × 48 cores × 2 threads/core = **192 vCPU**(96 實體核心) | | RAM | 2.0 TiB | | 虛擬化 | KVM(完整虛擬化) | | OpenVINO | CPU only,2026.4.0(build `2026.4.0-22959-99c81491cc3-releases/2026/4`) | | 設定 | `num_inference_steps=40`、`true_cfg_scale=1.0`、1024×1024 | **總結** | 指標 | 數值 | | --- | --- | | 解析度 | 1024×1024 | | 平均總耗時 | **869.80 s / 張** | | 平均單步耗時 | 21.57 s | | 最快 / 最慢 | 798.94 s / 980.80 s | | 總計(5 張) | 4349.00 s | | 記憶體高水位 | 40,211 MB | - **512px 沒有獨立耗時資料**:`outputs/*_512.png` 是 1024px 輸出的縮圖,非重新推理。 - 文字編碼與 VAE decode 的時間**已包含**在總耗時內。 **逐張結果(1024×1024)** | # | Prompt | Seed | 總耗時 (s) | 平均單步 (s) | true_cfg_scale | | --- | --- | --- | --- | --- | --- | | 01_hanfu | 42 | 980.80 | 24.257 | 1.0 | | 02_astronaut | 43 | 865.67 | 21.497 | 1.0 | | 03_taipei | 44 | 871.46 | 21.645 | 1.0 | | 04_shiba | 45 | 832.12 | 20.592 | 1.0 | | 05_ink | 46 | 798.94 | 19.846 | 1.0 | | | **平均** | **869.80** | **21.568** | | **模型大小** 以下為 repo 內 `openvino_model.bin` 的**實際位元組數**(Git LFS 記錄值)。 | 元件 | 位元組 | 大小 | 精度 | | --- | ---: | ---: | --- | | `text_encoder` | 4,256,132,893 | 4.26 GB | INT4 | | `text_encoder_i2i` | 4,256,132,733 | 4.26 GB | INT4 | | `transformer` | 3,696,679,634 | 3.70 GB | INT4 | | `vision_encoder` | 577,780,620 | 0.58 GB | — | | `vae_decoder` | 253,235,856 | 0.25 GB | — | | `vae_encoder` | 78,002,698 | 0.08 GB | — | | **合計** | **13,117,964,434** | **13.12 GB** | | FP16 匯出模型約 **43.6 GB**(transformer 14.23 GB + text_encoder 15.14 GB + text_encoder_i2i 15.14 GB + vision_encoder 1.1 GB + vae_decoder 484 MB + vae_encoder 150 MB)——取自原始轉換紀錄。 INT8 部分(vision encoder + VAE)為 **0.88 GB**,佔 INT4 總量的 7%。 `text_encoder` 與 `text_encoder_i2i` 各自 INT4 後仍有 4.26 GB,是整體壓縮率(~70%)的主要限制來源。 ## 檔案結構 ```text . ├── README.md # 本文件 ├── REPORT.md # 完整轉換 + 實測報告 ├── model_index.json # diffusers pipeline 索引(QwenImage21Pipeline) ├── openvino_config.json # OpenVINO 量化設定 ├── inference_int4.py # 單張推論範例 ├── generate5.py # 5 組 prompt 批次生成 + benchmark ├── quantize_int4.py # FP16 OV → INT4 OV 量化腳本 ├── transformer/ # INT4 QwenImage21Transformer2DModel(32 層) ├── text_encoder/ # INT4 Qwen3VLForConditionalGeneration ├── text_encoder_i2i/ # INT4 第二 text encoder(編輯用) ├── vision_encoder/ # INT8 Qwen3VLVisionModel ├── processor/ # Qwen3VLProcessor + tokenizer ├── vae_encoder/ # INT8 VAE encoder ├── vae_decoder/ # INT8 VAE decoder ├── scheduler/ # FlowMatchEulerDiscreteScheduler 設定 ├── examples/ # 5 張展示圖(與 outputs 1024 相同) └── outputs/ # 10 張實測圖 + 3 個資料檔 ├── *_1024.png # 主測組(5 張) ├── *_512.png # 對照組(5 張,為 1024 縮圖) ├── benchmark.json # 逐 step 耗時 + 記憶體 + 系統資訊 ├── benchmark_quantization.json # 量化設定與耗時 └── prompts.txt # 5 組 prompt 與 seed ``` ## 從零復現 ```bash # 1. 匯出 FP16 OpenVINO 模型 optimum-cli export openvino \ -m Qwen/Qwen-Image-2.1 \ --task text-to-image \ --library diffusers \ --weight-format fp16 \ ./Qwen-Image-2.1-ov-fp16 # 2. NNCF weight-only INT4 量化 python quantize_int4.py --fp16-dir ./Qwen-Image-2.1-ov-fp16 \ --int4-dir ./Qwen-Image-2.1-ov-int4 # 3. 單張推論 python inference_int4.py # 4. 批次生成 5 組 + benchmark python generate5.py --outdir outputs ``` 量化設定: ```python from optimum.intel.openvino.configuration import ( OVConfig, OVWeightQuantizationConfig, OVPipelineQuantizationConfig, ) int4_config = OVWeightQuantizationConfig( bits=4, sym=False, group_size=128, group_size_fallback="adjust", ratio=1.0, ) pipeline_config = OVPipelineQuantizationConfig( quantization_configs={ "transformer": int4_config, "text_encoder": int4_config, "text_encoder_i2i": int4_config, }, default_config=OVWeightQuantizationConfig(bits=8), ) ``` ## 已知限制 - **CPU 速度最慢**:1024×1024 / 40 steps 平均 **870 s / 張**(約 14.5 分鐘),是這批 repo 中最慢的,純 CPU 實務上不太實用。 - **記憶體需求最高**:INT4 模型 13.12 GB,實測行程結束時 RSS 達 **40.2 GB**,建議至少預留 48 GB 可用記憶體。 - **需要 diffusers 主分支**:`diffusers==0.41.0.dev0`(含 `QwenImage21Pipeline`,v0.40.0 之後才加入)。`optimum-intel` 需使用 `2.3.0.dev0+32a317a`。 - **需要對上游套件套用兩處修改**: - 1. `transformers` 的相依套件上限(hub cap `<1.0` → `<2.0`) - 2. `optimum-intel` 的 `modeling_visual_language.py`(transformers ≥ 5 的 `VisionRotaryEmbedding` 別名) - **512px 對照為縮圖**:`outputs/*_512.png` 是 1024px 輸出的 LANCZOS 縮圖,非重新推理。 - **`model_index.json` 未列出 `text_encoder_i2i` 與 `vision_encoder`**:這兩個目錄存在於 repo 中但未在索引中宣告,屬 optimum-intel 匯出的已知落差。 - **10 張風格展示圖無實測紀錄**:只有 5 張主測圖有 benchmark 資料。 ## 授權與出處 - **來源模型**:[`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1) - **授權**:`qwen-research`(`license: other`) - **完整條款**:[LICENSE](https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE) - **轉換**:僅做格式轉換與權重量化,模型權重來自來源模型 使用本模型時請遵守來源模型的授權條款。 ---
**Made with OpenVINO + optimum-intel + NNCF**