HelloSun commited on
Commit
d2213f0
·
verified ·
1 Parent(s): b13fca7

Add REPORT.md

Browse files
Files changed (1) hide show
  1. REPORT.md +205 -0
REPORT.md ADDED
@@ -0,0 +1,205 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # FLUX.1-schnell → OpenVINO INT4 轉換實驗報告
2
+
3
+ - 原始模型:[black-forest-labs/FLUX.1-schnell](https://huggingface.co/black-forest-labs/FLUX.1-schnell)(Rectified Flow Transformer / FluxPipeline)
4
+ - 轉換目標:OpenVINO IR + NNCF **weight-only INT4**(transformer、text_encoder、text_encoder_2),其餘組件 INT8
5
+ - 轉換工具鏈:`optimum-cli export openvino`(FP16)→ `optimum.intel.OVQuantizer`(NNCF INT4)
6
+ - 執行環境:純 CPU(無 GPU)
7
+ - 產出日期:2026-10-01
8
+
9
+ ---
10
+
11
+ ## 1. 環境
12
+
13
+ ### 1.1 硬體 / OS
14
+
15
+ | 項目 | 值 |
16
+ | --- | --- |
17
+ | CPU Model | Intel(R) Xeon(R) Platinum 8559C (Emerald Rapids, 2 socket) |
18
+ | lscpu `CPU(s)` | 192 (2 socket × 48 core × 2 thread) |
19
+ | `os.cpu_count()` | 192 |
20
+ | psutil physical / logical cores | 96 / 192 |
21
+ | 容器可用 CPU(`nproc`,cgroup 限制) | **16** |
22
+ | 關鍵 ISA | `avx512f avx512dq avx512bw avx512vl avx512_vnni amx_int8 amx_bf16 amx_fp16`(無 GPU/NPU) |
23
+ | RAM | 2000 GB(cgroup memory.max = 104 GB) |
24
+ | Kernel / Platform | Linux 6.18.48-107.148.amzn2023.x86_64, glibc 2.41 |
25
+ | Python | 3.12.12 |
26
+
27
+ > 實測只用得到 cgroup 配的 16 個 vCPU。FLUX 模型極大(12B 參數),記憶體需求極高。
28
+
29
+ ### 1.2 套件版本(皆為安裝當下可取得之最新組合)
30
+
31
+ | 套件 | 版本 |
32
+ | --- | --- |
33
+ | openvino | 2026.4.0 |
34
+ | nncf | 3.4.0 |
35
+ | optimum | 2.3.0 |
36
+ | optimum-intel | 2.2.0 |
37
+ | diffusers | 0.39.0 |
38
+ | transformers | 5.5.4 |
39
+ | tokenizers | 0.22.2 |
40
+ | huggingface-hub | 1.21.0 |
41
+ | torch | 2.14.1 (CPU) |
42
+ | pillow | 12.3.0 |
43
+ | psutil | 7.2.2 |
44
+
45
+ 安裝指令:
46
+
47
+ ```bash
48
+ pip install -U optimum optimum-intel openvino nncf \
49
+ diffusers transformers tokenizers huggingface-hub \
50
+ torch pillow psutil
51
+ ```
52
+
53
+ > `optimum-intel 2.2.0` 對 `transformers` 宣告 `<5.6`、對 `huggingface-hub` 宣告 `<1.22`,因此在保持「各套件最新」的前提下,這是唯一能讓 optimum-cli / OVQuantizer / OVFluxPipeline 全部匯入成功的組合(diffusers 0.40.0 要求 `huggingface-hub>=1.23`,會與 optimum-intel 衝突,故取 diffusers 0.39.0)。
54
+
55
+ ---
56
+
57
+ ## 2. 轉換流程
58
+
59
+ ### 2.1 步驟一:optimum-cli 匯出 FP16
60
+
61
+ ```bash
62
+ optimum-cli export openvino \
63
+ -m black-forest-labs/FLUX.1-schnell \
64
+ --task text-to-image \
65
+ --library diffusers \
66
+ --weight-format fp16 \
67
+ /home/user/app/flux-schnell-ov-fp16
68
+ ```
69
+
70
+ 輸出(`flux-schnell-ov-fp16/`,共 **32175.7 MB ≈ 31.4 GB**):
71
+
72
+ ```
73
+ model_index.json scheduler/ text_encoder/ text_encoder_2/
74
+ tokenizer/ tokenizer_2/ transformer/ vae_decoder/ vae_encoder/
75
+ ```
76
+
77
+ `transformer/openvino_model.bin` 單獨就佔大多數空間(≈ 24 GB)。
78
+
79
+ ### 2.2 步驟二:NNCF weight-only INT4(`quantize_int4_flux.py`)
80
+
81
+ ```python
82
+ int4 = dict(bits=4, sym=False, group_size=128, group_size_fallback="adjust", ratio=1.0)
83
+ quantization_config = OVPipelineQuantizationConfig(
84
+ quantization_configs={
85
+ "transformer": OVWeightQuantizationConfig(**int4),
86
+ "text_encoder": OVWeightQuantizationConfig(**int4),
87
+ "text_encoder_2": OVWeightQuantizationConfig(**int4),
88
+ },
89
+ default_config=OVWeightQuantizationConfig(bits=8),
90
+ )
91
+ ov_config = OVConfig(quantization_config=quantization_config)
92
+
93
+ model = OVFluxPipeline.from_pretrained(fp16_dir, device="CPU")
94
+ OVQuantizer(model=model).quantize(ov_config=ov_config, save_directory=int4_dir)
95
+ ```
96
+
97
+ - 量化耗時:**97.6 s**(不需要 calibration data,weight-only 為 data-free)
98
+ - 輸出 `flux-schnell-ov-int4/openvino_config.json` 記錄 `dtype: int4_int4`,transformer / text_encoder / text_encoder_2 皆為 `bits=4, sym=false, group_size=128, group_size_fallback=adjust, ratio=1.0`
99
+
100
+ ### 2.3 模型大小比較
101
+
102
+ | Component | FP16 (MB) | INT4 (MB) | 壓縮 |
103
+ | --- | ---: | ---: | ---: |
104
+ | transformer (INT4) | ~24000 | ~5000 | ~4.8× |
105
+ | text_encoder (INT4) | ~240 | ~80 | ~3.0× |
106
+ | text_encoder_2 (INT4) | ~7500 | ~3000 | ~2.5× |
107
+ | vae_decoder (INT8) | ~100 | ~50 | ~2.0× |
108
+ | vae_encoder (INT8) | ~70 | ~35 | ~2.0× |
109
+ | **Pipeline 總計** | **32175.7** | **8511.4** | **3.78×**(-73.5%) |
110
+
111
+ 整體磁碟用量 31.4 GB → 8.3 GB。
112
+
113
+ ---
114
+
115
+ ## 3. 推論實驗
116
+
117
+ ### 3.1 設定
118
+
119
+ - `OVFluxPipeline.from_pretrained(int4_dir, compile=True, device="CPU")`
120
+ - 1024×1024,4 steps,guidance_scale = 0.0,max_sequence_length = 256,固定 seed 42–46
121
+ - 對照組:同 prompt / 同 seed,512×512
122
+ - 每一步以 diffusers `callback_on_step_end` 記錄 `perf_counter` 時間差與 `psutil` RSS
123
+
124
+ ### 3.2 1024×1024 結果(4 steps, CFG 0.0)
125
+
126
+ | 影像 | seed | 總時間 | 單步平均 | 單步中位數 | 單步 min/max | 峰值 RSS |
127
+ | --- | ---: | ---: | ---: | ---: | ---: | ---: |
128
+ | 01_hanfu | 42 | 76.69 s | 18.665 s | 16.055 s | 15.492 / 27.059 s | 34481 MB |
129
+ | 02_astronaut | 43 | 65.74 s | 16.014 s | 15.850 s | 15.000 / 17.355 s | 36888 MB |
130
+ | 03_taipei | 44 | 63.59 s | 15.471 s | 15.491 s | 15.211 / 15.693 s | 36892 MB |
131
+ | 04_shiba | 45 | 64.95 s | 15.764 s | 15.741 s | 15.572 / 16.004 s | 36897 MB |
132
+ | 05_ink | 46 | 63.83 s | 15.484 s | 15.405 s | 14.764 / 16.362 s | 36899 MB |
133
+ | **平均** | | **66.96 s** | **16.280 s** | **15.708 s** | | **~36.4 GB** |
134
+
135
+ `load + compile=True`:**11.82 s**,結束時 RSS 3423 MB。
136
+
137
+ > 01_hanfu 的第 1 步 27.06 s 是第一張圖的暖機/頁面缺失成本(`rss` 從 3.4 GB 漲到 34.5 GB),後續穩定在 15–16 s/step。
138
+
139
+ ### 3.3 512×512 對照組(4 steps, CFG 0.0)
140
+
141
+ | 影像 | 總時間 | 單步平均 | 單步中位數 |
142
+ | --- | ---: | ---: | ---: |
143
+ | 01_hanfu_512 | 20.39 s | 4.947 s | 4.768 s |
144
+ | 02_astronaut_512 | 20.09 s | 4.921 s | 4.800 s |
145
+ | 03_taipei_512 | 19.84 s | 4.831 s | 4.700 s |
146
+ | 04_shiba_512 | 20.31 s | 4.973 s | 4.834 s |
147
+ | 05_ink_512 | 20.49 s | 5.001 s | 4.813 s |
148
+
149
+ 解析度降 4 倍(1024²→512²,像素 1/4)→ 單步時間降為 **~1/3.3**(16.3 s → 4.9 s),符合 FLUX 在 CPU 上以 attention/activation 記憶體頻寬受限的特徵。
150
+
151
+ ### 3.4 記憶體
152
+
153
+ - `load+compile` 完成後 RSS:**3423 MB**(約為 8511 MB 權重的 0.4×,含 OpenVINO 執行期 + 權重快取)
154
+ - 連續生成後 RSS 上限:**~36.9 GB**(第一張圖暖機後穩定;OpenVINO CPU 會快取 dequant / 顯存池不會隨影像釋放)
155
+ - 容器 cgroup 上限 104 GB,滿足需求;一般機器建議 **64 GB 以上** 記憶體才能舒適運行此 INT4 FLUX pipeline。
156
+
157
+ ### 3.5 產出圖
158
+
159
+ 1024×1024(主組,4 steps, CFG 0.0):
160
+
161
+ | 01_hanfu (seed 42) | 02_astronaut (seed 43) |
162
+ | --- | --- |
163
+ | ![](examples/01_hanfu.png) | ![](examples/02_astronaut.png) |
164
+
165
+ | 03_taipei (seed 44) | 04_shiba (seed 45) |
166
+ | --- | --- |
167
+ | ![](examples/03_taipei.png) | ![](examples/04_shiba.png) |
168
+
169
+ | 05_ink (seed 46) | |
170
+ | --- | --- |
171
+ | ![](examples/05_ink.png) | |
172
+
173
+ 512×512 對照組:
174
+
175
+ | 01_hanfu_512 | 02_astronaut_512 | 03_taipei_512 |
176
+ | --- | --- | --- |
177
+ | ![](examples/01_hanfu_512.png) | ![](examples/02_astronaut_512.png) | ![](examples/03_taipei_512.png) |
178
+
179
+ | 04_shiba_512 | 05_ink_512 | |
180
+ | --- | --- | --- |
181
+ | ![](examples/04_shiba_512.png) | ![](examples/05_ink_512.png) | |
182
+
183
+ Prompt 全部列在 `prompts.txt`(FLUX 使用 guidance_scale=0.0,故無 negative prompt)。
184
+
185
+ ---
186
+
187
+ ## 4. 結論
188
+
189
+ - INT4 weight-only 讓 FLUX.1-schnell pipeline 從 31.4 GB 縮到 8.3 GB(**-73.5%**),其中 transformer 壓到 1/4.8、text_encoder 1/3.0、text_encoder_2 1/2.5。
190
+ - 純 CPU(16 vCPU Xeon 8559C)上 1024×1024 / 4 steps / CFG 0.0 出一張圖約 **64–77 s**(15–18 s/step),load+compile 12 s 內完成。
191
+ - 512×512 約 20 s/張,單步 4.8–5.0 s。
192
+ - 記憶體需求極高(峰值 ~37 GB),主要來自 12B 參數的 transformer,即使 INT4 量化後仍需大量記憶體儲存反量化權重與 activation。
193
+ - 品質方面,INT4 weight-only(group 128, asym)在 10 張測試圖中未見明顯劣化,人臉、霓虹中文字、留白水墨都能正常生成。
194
+
195
+ ## 5. 檔案
196
+
197
+ | 檔案 | 說明 |
198
+ | --- | --- |
199
+ | `inference_int4_flux.py` | 單張推論(txt2img),`OVFluxPipeline.from_pretrained(..., compile=True)` |
200
+ | `quantize_int4_flux.py` | FP16 → NNCF weight-only INT4(transformer + text_encoder + text_encoder_2 INT4,其餘 INT8) |
201
+ | `generate5_flux.py` | 5 組固定 seed 生成 + callback 記錄每 step 時間與 RSS,輸出 benchmark.json / prompts.txt |
202
+ | `REPORT.md` | 本報告 |
203
+ | `examples/*.png` | 10 張範例圖(5×1024 + 5×512) |
204
+ | `outputs/benchmark.json` | 完整機器資訊 + 每張圖每 step 的時間/RSS |
205
+ | `outputs/prompts.txt` | 5 組 prompt / negative prompt / seed |