HelloSun commited on
Commit
5267c71
·
verified ·
1 Parent(s): 4825e8b

Add REPORT.md

Browse files
Files changed (1) hide show
  1. REPORT.md +147 -0
REPORT.md ADDED
@@ -0,0 +1,147 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Qwen-Image-2.1 OpenVINO INT4 實驗報告 (REPORT.md)
2
+
3
+ - 原始模型: `Qwen/Qwen-Image-2.1` (QwenImage21Pipeline, 7B visual generation, 32 Single-Stream DiT layers)
4
+ - 轉換目標: OpenVINO FP16 -> NNCF weight-only INT4
5
+ - 目標 repo: `HelloSun/Qwen-Image-2.1-OpenVINO-INT4`
6
+ - 時間: 2026-09-29 UTC
7
+ - 原模型官方範例參數 (https://huggingface.co/Qwen/Qwen-Image-2.1 `QwenImage21Pipeline`):
8
+ ```python
9
+ image = pipe(prompt=..., width=2048, height=2048, num_inference_steps=40,
10
+ generator=torch.Generator("cuda").manual_seed(42)).images[0]
11
+ ```
12
+ `num_inference_steps=40`, `true_cfg_scale=1.0` (預設, 無 guidance), 官方 1:1 為 2048x2048。
13
+ 本實驗依任務要求固定解析度 `1024x1024`, steps 採用官方 `40`, `true_cfg_scale=1.0`。
14
+
15
+ ## 1. 環境 (最新版)
16
+
17
+ 任務要求「環境盡量用新版本 diffusers transformers tokenizers huggingface-hub optimum-intel optimum openvino nncf torch pillow psutil,不可用其他版本組合」。
18
+ Qwen-Image-2.1 的 `QwenImage21Pipeline` 是 diffusers commit 6256aa766 "Add Qwen-Image 2.1 (#14804)" (post-v0.40.0) 才加入,
19
+ pinned `diffusers==0.37.1` 無法載入 (`AttributeError`),故採用以下最新版 (實測 `pip list`):
20
+
21
+ ```
22
+ diffusers==0.41.0.dev0 (git main, 含 QwenImage21)
23
+ transformers==5.10.4
24
+ tokenizers==0.22.2
25
+ huggingface-hub==1.33.0
26
+ optimum==2.3.0
27
+ optimum-intel==2.3.0.dev0+32a317a (main, 首個支援 QwenImage21 的版本)
28
+ openvino==2026.4.0 (build 2026.4.0-22959-99c81491cc3-releases/2026/4)
29
+ openvino-telemetry==2025.2.0, openvino-tokenizers==2026.4.0.0
30
+ nncf==3.4.0 torch==2.14.0 pillow==12.3.0 psutil==7.2.2
31
+ ```
32
+
33
+ 另含 `transformers` dependency table patch (hub cap `<1.0` -> `<2.0`) 與
34
+ `optimum-intel modeling_visual_language.py` patch (transformers>=5 `VisionRotaryEmbedding` alias),
35
+ 僅為相容性修正,推理未用到該路徑。
36
+
37
+ ## 2. CPU 真實能力
38
+
39
+ - `os.cpu_count()` / `psutil.cpu_count(logical=True)`: **192**
40
+ - `psutil.cpu_count(logical=False)`: **96**
41
+ - `lscpu`:
42
+ ```
43
+ Architecture: x86_64
44
+ CPU(s): 192 (On-line 0-191)
45
+ Vendor ID: GenuineIntel
46
+ Model name: Intel(R) Xeon(R) Platinum 8559C
47
+ CPU family 6, Model 207, Stepping 2
48
+ Thread(s) per core: 2, Core(s) per socket: 48, Socket(s): 2
49
+ BogoMIPS: 4800.00
50
+ L1d 4.5 MiB (96 instances), L1i 3 MiB (96), L2 192 MiB (96), L3 640 MiB (2)
51
+ NUMA node(s): 2 (node0: 0-47,96-143; node1: 48-95,144-191)
52
+ Hypervisor: KVM, full virtualization
53
+ Flags含 AVX2/AVX512F/AVX512_BF16/AVX512_FP16/AMX_BF16/AMX_INT8 (有利 OpenVINO CPU 加速)
54
+ Mem: ~2.0TiB total (實驗時 available ~840Gi, free ~307Gi), Swap 1.6TiB
55
+ ```
56
+ - OpenVINO 版本: **2026.4.0-22959-99c81491cc3-releases/2026/4**
57
+ - 完整 `lscpu` 見 `outputs/benchmark.json` (`cpu.lscpu` 欄位)。
58
+
59
+ ## 3. 轉換流程
60
+
61
+ ### 3.1 FP16 導出 (optimum-cli)
62
+ 任務範例:
63
+ ```
64
+ optimum-cli export openvino -m Tongyi-MAI/Z-Image-Turbo --task text-to-image --library diffusers --weight-format fp16 /home/user/app/z-image-turbo-ov-fp16
65
+ ```
66
+ 本模型等價指令:
67
+ ```
68
+ optimum-cli export openvino -m Qwen/Qwen-Image-2.1 --task text-to-image --library diffusers --weight-format fp16 /home/user/app/qwen-image-2.1-ov-fp16
69
+ ```
70
+ 實作以 `OVDiffusionPipeline.from_pretrained(MODEL_ID, export=True, compile=False, weight_format="fp16")` 執行 (底層同 optimum export 邏輯),已存在則跳過。本次 FP16 已存在,export_time=0s。
71
+ FP16 大小 (openvino_model.bin):
72
+ - transformer `14230249906` bytes (~13.25GiB), text_encoder `15136803293` bytes (~14.10GiB), text_encoder_i2i `15136803261` bytes (~14.10GiB)
73
+ - vae_decoder 484M, vae_encoder 150M, vision_encoder 1.1G; 整目錄約 44G。
74
+
75
+ ### 3.2 INT4 量化 (OVQuantizer + OVPipelineQuantizationConfig, NNCF weight-only)
76
+ `quantize_int4.py`:
77
+ ```python
78
+ from optimum.intel.openvino import OVQuantizer, OVConfig, OVPipelineQuantizationConfig, OVWeightQuantizationConfig
79
+ int4_cfg = OVWeightQuantizationConfig(bits=4, sym=False, group_size=128, group_size_fallback="adjust", ratio=1.0)
80
+ int8_default = OVWeightQuantizationConfig() # bits=8 預設 INT8
81
+ ov_config = OVConfig(quantization_config=OVPipelineQuantizationConfig(
82
+ quantization_configs={"transformer": int4_cfg, "text_encoder": int4_cfg, "text_encoder_i2i": int4_cfg},
83
+ default_config=int8_default))
84
+ quantizer = OVQuantizer.from_pretrained(OVDiffusionPipeline.from_pretrained(FP16_DIR, compile=False))
85
+ quantizer.quantize(save_directory=INT4_DIR, ov_config=ov_config)
86
+ ```
87
+ - `text_encoder_i2i` 為 Qwen-Image-2.1 editing 用第二 text encoder (同 Qwen3VL 架構),一併 INT4;其餘 (vae_decoder/vae_encoder/vision_encoder) 預設 INT8。
88
+ - 以 `ov_config=OVConfig(quantization_config=...)` 傳入,符合任務要求。
89
+ - 量化時間: **115.44s** (`outputs/benchmark_quantization.json`)。
90
+ - INT4 大小: transformer `3696679634` bytes (~3.44GiB), text_encoder `4256132893` bytes (~3.96GiB), text_encoder_i2i `4256132733` bytes (~3.96GiB), vae_decoder 243M, vae_encoder 76M, vision_encoder 553M; 整目錄約 13G。壓縮比約 3.4x。
91
+
92
+ ## 4. 推理配置
93
+
94
+ ```python
95
+ pipeline = OVDiffusionPipeline.from_pretrained(INT4_DIR, compile=True)
96
+ image = pipeline(prompt=..., num_inference_steps=40, height=1024, width=1024,
97
+ generator=torch.Generator().manual_seed(seed),
98
+ callback_on_step_end=StepCallback(),
99
+ callback_on_step_end_tensor_inputs=["latents"]).images[0]
100
+ ```
101
+ - `true_cfg_scale` 用預設 1.0 (官方範例未傳,即無 guidance,與任務 `guidance_scale 依照原模型建議` 一致)。
102
+ - `callback_on_step_end(pipe, step, timestep, callback_kwargs)` 記錄每 step wall time + `psutil.Process().memory_info().rss`。
103
+ - 5 組固定 prompt/seed 存到 `/home/user/app/outputs/` : `{name}_1024.png` + 512px 對照組 `{name}_512.png` (PIL resize 512x512,非重新推理)。
104
+ - 腳本: `inference_int4.py` (單張範例), `generate5.py` (5張+benchmark), `quantize_int4.py` (導出+量化)。
105
+
106
+ ## 5. 實驗數據
107
+
108
+ load+compile (INT4 `from_pretrained(compile=True)`): **14.69s**
109
+ final process RSS: **40210.9 MB** (~39.3GiB)
110
+
111
+ | name | seed | steps | true_cfg_scale | 解析度 | 總生圖時間(s) | 平均單步(s) | peak RSS(MB) | avg RSS(MB) |
112
+ |---|---|---|---|---|---|---|---|---|
113
+ | 01_hanfu | 42 | 40 | 1.0 | 1024x1024 | 980.80 | 24.26 | 36226.2 | 36225.9 |
114
+ | 02_astronaut | 43 | 40 | 1.0 | 1024x1024 | 865.67 | 21.50 | 39979.2 | 39979.2 |
115
+ | 03_taipei | 44 | 40 | 1.0 | 1024x1024 | 871.46 | 21.65 | 40185.4 | 40179.5 |
116
+ | 04_shiba | 45 | 40 | 1.0 | 1024x1024 | 832.12 | 20.59 | 40221.8 | 40221.8 |
117
+ | 05_ink | 46 | 40 | 1.0 | 1024x1024 | 798.94 | 19.85 | 40163.0 | 40163.0 |
118
+
119
+ 註: callback 在 40 steps 下觸發 39 次紀錄 (首步作為基準,`step_times` 長度 39);總生圖時間含首步 text-encoder/首步 overhead + VAE decode。
120
+ 記憶體: 首張 36.2GB,後續穩定 ~40GB (含 compiled model + peak activations, 1024x1024, 40 steps)。
121
+ 512px 對照組由 1024 圖直接 resize,非重新推理。
122
+
123
+ prompt 全文見 `outputs/prompts.txt`:
124
+ - 01_hanfu seed42: Young Chinese woman in red Hanfu ... photorealistic, ultra detailed, 8k
125
+ - 02_astronaut seed43: Astronaut in a jungle, cold color palette ...
126
+ - 03_taipei seed44: Cyberpunk street in Taipei at night ... 'TAIPEI' '台北' ...
127
+ - 04_shiba seed45: Cute Shiba Inu wearing a tiny astronaut helmet ...
128
+ - 05_ink seed46: Traditional Chinese ink wash landscape ...
129
+
130
+ 每步詳細時間與 RSS 見 `outputs/benchmark.json` (`results[].step_times`, `memory_usage_mb`)。
131
+
132
+ ## 6. 檔案清單
133
+
134
+ - `/home/user/app/qwen-image-2.1-ov-int4/` : INT4 模型 (已上傳至 HF repo 根目錄)
135
+ - `/home/user/app/outputs/benchmark.json` : 本報告機器可讀版 (含 lscpu/cpu/openvino/load/每步時間/RSS)
136
+ - `/home/user/app/outputs/prompts.txt` : 5 組 prompt+seed
137
+ - `/home/user/app/outputs/benchmark_quantization.json` : export/quantize 時間
138
+ - `/home/user/app/outputs/*_1024.png` + `*_512.png` : 生成圖
139
+ - `/home/user/app/examples/*.png` : 同 1024 圖,供 README 展示
140
+ - `/home/user/app/quantize_int4.py`, `generate5.py`, `inference_int4.py`
141
+ - `/home/user/app/REPORT.md` (本檔), HF README 見 repo `README.md`
142
+
143
+ ## 7. 注意事項
144
+
145
+ - VAE `scaling_factor missing` warning 為 optimum-intel + diffusers 已知提示,不影響生成 (VAE 仍以 INT8 量化輸出正常 PNG)。
146
+ - `optimum` 套件存在多 distribution 警告 (`Multiple distributions found for package optimum`) 為環境預裝特性,不影響功能。
147
+ - CPU 推理 1024x1024 40 steps 約 800-980s/張,單步約 19.8-24.3s;若需更快可降 steps/解析度或使用 GPU。