Add REPORT.md
Browse files
REPORT.md
ADDED
|
@@ -0,0 +1,147 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Qwen-Image-2.1 OpenVINO INT4 實驗報告 (REPORT.md)
|
| 2 |
+
|
| 3 |
+
- 原始模型: `Qwen/Qwen-Image-2.1` (QwenImage21Pipeline, 7B visual generation, 32 Single-Stream DiT layers)
|
| 4 |
+
- 轉換目標: OpenVINO FP16 -> NNCF weight-only INT4
|
| 5 |
+
- 目標 repo: `HelloSun/Qwen-Image-2.1-OpenVINO-INT4`
|
| 6 |
+
- 時間: 2026-09-29 UTC
|
| 7 |
+
- 原模型官方範例參數 (https://huggingface.co/Qwen/Qwen-Image-2.1 `QwenImage21Pipeline`):
|
| 8 |
+
```python
|
| 9 |
+
image = pipe(prompt=..., width=2048, height=2048, num_inference_steps=40,
|
| 10 |
+
generator=torch.Generator("cuda").manual_seed(42)).images[0]
|
| 11 |
+
```
|
| 12 |
+
`num_inference_steps=40`, `true_cfg_scale=1.0` (預設, 無 guidance), 官方 1:1 為 2048x2048。
|
| 13 |
+
本實驗依任務要求固定解析度 `1024x1024`, steps 採用官方 `40`, `true_cfg_scale=1.0`。
|
| 14 |
+
|
| 15 |
+
## 1. 環境 (最新版)
|
| 16 |
+
|
| 17 |
+
任務要求「環境盡量用新版本 diffusers transformers tokenizers huggingface-hub optimum-intel optimum openvino nncf torch pillow psutil,不可用其他版本組合」。
|
| 18 |
+
Qwen-Image-2.1 的 `QwenImage21Pipeline` 是 diffusers commit 6256aa766 "Add Qwen-Image 2.1 (#14804)" (post-v0.40.0) 才加入,
|
| 19 |
+
pinned `diffusers==0.37.1` 無法載入 (`AttributeError`),故採用以下最新版 (實測 `pip list`):
|
| 20 |
+
|
| 21 |
+
```
|
| 22 |
+
diffusers==0.41.0.dev0 (git main, 含 QwenImage21)
|
| 23 |
+
transformers==5.10.4
|
| 24 |
+
tokenizers==0.22.2
|
| 25 |
+
huggingface-hub==1.33.0
|
| 26 |
+
optimum==2.3.0
|
| 27 |
+
optimum-intel==2.3.0.dev0+32a317a (main, 首個支援 QwenImage21 的版本)
|
| 28 |
+
openvino==2026.4.0 (build 2026.4.0-22959-99c81491cc3-releases/2026/4)
|
| 29 |
+
openvino-telemetry==2025.2.0, openvino-tokenizers==2026.4.0.0
|
| 30 |
+
nncf==3.4.0 torch==2.14.0 pillow==12.3.0 psutil==7.2.2
|
| 31 |
+
```
|
| 32 |
+
|
| 33 |
+
另含 `transformers` dependency table patch (hub cap `<1.0` -> `<2.0`) 與
|
| 34 |
+
`optimum-intel modeling_visual_language.py` patch (transformers>=5 `VisionRotaryEmbedding` alias),
|
| 35 |
+
僅為相容性修正,推理未用到該路徑。
|
| 36 |
+
|
| 37 |
+
## 2. CPU 真實能力
|
| 38 |
+
|
| 39 |
+
- `os.cpu_count()` / `psutil.cpu_count(logical=True)`: **192**
|
| 40 |
+
- `psutil.cpu_count(logical=False)`: **96**
|
| 41 |
+
- `lscpu`:
|
| 42 |
+
```
|
| 43 |
+
Architecture: x86_64
|
| 44 |
+
CPU(s): 192 (On-line 0-191)
|
| 45 |
+
Vendor ID: GenuineIntel
|
| 46 |
+
Model name: Intel(R) Xeon(R) Platinum 8559C
|
| 47 |
+
CPU family 6, Model 207, Stepping 2
|
| 48 |
+
Thread(s) per core: 2, Core(s) per socket: 48, Socket(s): 2
|
| 49 |
+
BogoMIPS: 4800.00
|
| 50 |
+
L1d 4.5 MiB (96 instances), L1i 3 MiB (96), L2 192 MiB (96), L3 640 MiB (2)
|
| 51 |
+
NUMA node(s): 2 (node0: 0-47,96-143; node1: 48-95,144-191)
|
| 52 |
+
Hypervisor: KVM, full virtualization
|
| 53 |
+
Flags含 AVX2/AVX512F/AVX512_BF16/AVX512_FP16/AMX_BF16/AMX_INT8 (有利 OpenVINO CPU 加速)
|
| 54 |
+
Mem: ~2.0TiB total (實驗時 available ~840Gi, free ~307Gi), Swap 1.6TiB
|
| 55 |
+
```
|
| 56 |
+
- OpenVINO 版本: **2026.4.0-22959-99c81491cc3-releases/2026/4**
|
| 57 |
+
- 完整 `lscpu` 見 `outputs/benchmark.json` (`cpu.lscpu` 欄位)。
|
| 58 |
+
|
| 59 |
+
## 3. 轉換流程
|
| 60 |
+
|
| 61 |
+
### 3.1 FP16 導出 (optimum-cli)
|
| 62 |
+
任務範例:
|
| 63 |
+
```
|
| 64 |
+
optimum-cli export openvino -m Tongyi-MAI/Z-Image-Turbo --task text-to-image --library diffusers --weight-format fp16 /home/user/app/z-image-turbo-ov-fp16
|
| 65 |
+
```
|
| 66 |
+
本模型等價指令:
|
| 67 |
+
```
|
| 68 |
+
optimum-cli export openvino -m Qwen/Qwen-Image-2.1 --task text-to-image --library diffusers --weight-format fp16 /home/user/app/qwen-image-2.1-ov-fp16
|
| 69 |
+
```
|
| 70 |
+
實作以 `OVDiffusionPipeline.from_pretrained(MODEL_ID, export=True, compile=False, weight_format="fp16")` 執行 (底層同 optimum export 邏輯),已存在則跳過。本次 FP16 已存在,export_time=0s。
|
| 71 |
+
FP16 大小 (openvino_model.bin):
|
| 72 |
+
- transformer `14230249906` bytes (~13.25GiB), text_encoder `15136803293` bytes (~14.10GiB), text_encoder_i2i `15136803261` bytes (~14.10GiB)
|
| 73 |
+
- vae_decoder 484M, vae_encoder 150M, vision_encoder 1.1G; 整目錄約 44G。
|
| 74 |
+
|
| 75 |
+
### 3.2 INT4 量化 (OVQuantizer + OVPipelineQuantizationConfig, NNCF weight-only)
|
| 76 |
+
`quantize_int4.py`:
|
| 77 |
+
```python
|
| 78 |
+
from optimum.intel.openvino import OVQuantizer, OVConfig, OVPipelineQuantizationConfig, OVWeightQuantizationConfig
|
| 79 |
+
int4_cfg = OVWeightQuantizationConfig(bits=4, sym=False, group_size=128, group_size_fallback="adjust", ratio=1.0)
|
| 80 |
+
int8_default = OVWeightQuantizationConfig() # bits=8 預設 INT8
|
| 81 |
+
ov_config = OVConfig(quantization_config=OVPipelineQuantizationConfig(
|
| 82 |
+
quantization_configs={"transformer": int4_cfg, "text_encoder": int4_cfg, "text_encoder_i2i": int4_cfg},
|
| 83 |
+
default_config=int8_default))
|
| 84 |
+
quantizer = OVQuantizer.from_pretrained(OVDiffusionPipeline.from_pretrained(FP16_DIR, compile=False))
|
| 85 |
+
quantizer.quantize(save_directory=INT4_DIR, ov_config=ov_config)
|
| 86 |
+
```
|
| 87 |
+
- `text_encoder_i2i` 為 Qwen-Image-2.1 editing 用第二 text encoder (同 Qwen3VL 架構),一併 INT4;其餘 (vae_decoder/vae_encoder/vision_encoder) 預設 INT8。
|
| 88 |
+
- 以 `ov_config=OVConfig(quantization_config=...)` 傳入,符合任務要求。
|
| 89 |
+
- 量化時間: **115.44s** (`outputs/benchmark_quantization.json`)。
|
| 90 |
+
- INT4 大小: transformer `3696679634` bytes (~3.44GiB), text_encoder `4256132893` bytes (~3.96GiB), text_encoder_i2i `4256132733` bytes (~3.96GiB), vae_decoder 243M, vae_encoder 76M, vision_encoder 553M; 整目錄約 13G。壓縮比約 3.4x。
|
| 91 |
+
|
| 92 |
+
## 4. 推理配置
|
| 93 |
+
|
| 94 |
+
```python
|
| 95 |
+
pipeline = OVDiffusionPipeline.from_pretrained(INT4_DIR, compile=True)
|
| 96 |
+
image = pipeline(prompt=..., num_inference_steps=40, height=1024, width=1024,
|
| 97 |
+
generator=torch.Generator().manual_seed(seed),
|
| 98 |
+
callback_on_step_end=StepCallback(),
|
| 99 |
+
callback_on_step_end_tensor_inputs=["latents"]).images[0]
|
| 100 |
+
```
|
| 101 |
+
- `true_cfg_scale` 用預設 1.0 (官方範例未傳,即無 guidance,與任務 `guidance_scale 依照原模型建議` 一致)。
|
| 102 |
+
- `callback_on_step_end(pipe, step, timestep, callback_kwargs)` 記錄每 step wall time + `psutil.Process().memory_info().rss`。
|
| 103 |
+
- 5 組固定 prompt/seed 存到 `/home/user/app/outputs/` : `{name}_1024.png` + 512px 對照組 `{name}_512.png` (PIL resize 512x512,非重新推理)。
|
| 104 |
+
- 腳本: `inference_int4.py` (單張範例), `generate5.py` (5張+benchmark), `quantize_int4.py` (導出+量化)。
|
| 105 |
+
|
| 106 |
+
## 5. 實驗數據
|
| 107 |
+
|
| 108 |
+
load+compile (INT4 `from_pretrained(compile=True)`): **14.69s**
|
| 109 |
+
final process RSS: **40210.9 MB** (~39.3GiB)
|
| 110 |
+
|
| 111 |
+
| name | seed | steps | true_cfg_scale | 解析度 | 總生圖時間(s) | 平均單步(s) | peak RSS(MB) | avg RSS(MB) |
|
| 112 |
+
|---|---|---|---|---|---|---|---|---|
|
| 113 |
+
| 01_hanfu | 42 | 40 | 1.0 | 1024x1024 | 980.80 | 24.26 | 36226.2 | 36225.9 |
|
| 114 |
+
| 02_astronaut | 43 | 40 | 1.0 | 1024x1024 | 865.67 | 21.50 | 39979.2 | 39979.2 |
|
| 115 |
+
| 03_taipei | 44 | 40 | 1.0 | 1024x1024 | 871.46 | 21.65 | 40185.4 | 40179.5 |
|
| 116 |
+
| 04_shiba | 45 | 40 | 1.0 | 1024x1024 | 832.12 | 20.59 | 40221.8 | 40221.8 |
|
| 117 |
+
| 05_ink | 46 | 40 | 1.0 | 1024x1024 | 798.94 | 19.85 | 40163.0 | 40163.0 |
|
| 118 |
+
|
| 119 |
+
註: callback 在 40 steps 下觸發 39 次紀錄 (首步作為基準,`step_times` 長度 39);總生圖時間含首步 text-encoder/首步 overhead + VAE decode。
|
| 120 |
+
記憶體: 首張 36.2GB,後續穩定 ~40GB (含 compiled model + peak activations, 1024x1024, 40 steps)。
|
| 121 |
+
512px 對照組由 1024 圖直接 resize,非重新推理。
|
| 122 |
+
|
| 123 |
+
prompt 全文見 `outputs/prompts.txt`:
|
| 124 |
+
- 01_hanfu seed42: Young Chinese woman in red Hanfu ... photorealistic, ultra detailed, 8k
|
| 125 |
+
- 02_astronaut seed43: Astronaut in a jungle, cold color palette ...
|
| 126 |
+
- 03_taipei seed44: Cyberpunk street in Taipei at night ... 'TAIPEI' '台北' ...
|
| 127 |
+
- 04_shiba seed45: Cute Shiba Inu wearing a tiny astronaut helmet ...
|
| 128 |
+
- 05_ink seed46: Traditional Chinese ink wash landscape ...
|
| 129 |
+
|
| 130 |
+
每步詳細時間與 RSS 見 `outputs/benchmark.json` (`results[].step_times`, `memory_usage_mb`)。
|
| 131 |
+
|
| 132 |
+
## 6. 檔案清單
|
| 133 |
+
|
| 134 |
+
- `/home/user/app/qwen-image-2.1-ov-int4/` : INT4 模型 (已上傳至 HF repo 根目錄)
|
| 135 |
+
- `/home/user/app/outputs/benchmark.json` : 本報告機器可讀版 (含 lscpu/cpu/openvino/load/每步時間/RSS)
|
| 136 |
+
- `/home/user/app/outputs/prompts.txt` : 5 組 prompt+seed
|
| 137 |
+
- `/home/user/app/outputs/benchmark_quantization.json` : export/quantize 時間
|
| 138 |
+
- `/home/user/app/outputs/*_1024.png` + `*_512.png` : 生成圖
|
| 139 |
+
- `/home/user/app/examples/*.png` : 同 1024 圖,供 README 展示
|
| 140 |
+
- `/home/user/app/quantize_int4.py`, `generate5.py`, `inference_int4.py`
|
| 141 |
+
- `/home/user/app/REPORT.md` (本檔), HF README 見 repo `README.md`
|
| 142 |
+
|
| 143 |
+
## 7. 注意事項
|
| 144 |
+
|
| 145 |
+
- VAE `scaling_factor missing` warning 為 optimum-intel + diffusers 已知提示,不影響生成 (VAE 仍以 INT8 量化輸出正常 PNG)。
|
| 146 |
+
- `optimum` 套件存在多 distribution 警告 (`Multiple distributions found for package optimum`) 為環境預裝特性,不影響功能。
|
| 147 |
+
- CPU 推理 1024x1024 40 steps 約 800-980s/張,單步約 19.8-24.3s;若需更快可降 steps/解析度或使用 GPU。
|