File size: 8,541 Bytes
d2213f0 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 | # FLUX.1-schnell → OpenVINO INT4 轉換實驗報告
- 原始模型:[black-forest-labs/FLUX.1-schnell](https://huggingface.co/black-forest-labs/FLUX.1-schnell)(Rectified Flow Transformer / FluxPipeline)
- 轉換目標:OpenVINO IR + NNCF **weight-only INT4**(transformer、text_encoder、text_encoder_2),其餘組件 INT8
- 轉換工具鏈:`optimum-cli export openvino`(FP16)→ `optimum.intel.OVQuantizer`(NNCF INT4)
- 執行環境:純 CPU(無 GPU)
- 產出日期:2026-10-01
---
## 1. 環境
### 1.1 硬體 / OS
| 項目 | 值 |
| --- | --- |
| CPU Model | Intel(R) Xeon(R) Platinum 8559C (Emerald Rapids, 2 socket) |
| lscpu `CPU(s)` | 192 (2 socket × 48 core × 2 thread) |
| `os.cpu_count()` | 192 |
| psutil physical / logical cores | 96 / 192 |
| 容器可用 CPU(`nproc`,cgroup 限制) | **16** |
| 關鍵 ISA | `avx512f avx512dq avx512bw avx512vl avx512_vnni amx_int8 amx_bf16 amx_fp16`(無 GPU/NPU) |
| RAM | 2000 GB(cgroup memory.max = 104 GB) |
| Kernel / Platform | Linux 6.18.48-107.148.amzn2023.x86_64, glibc 2.41 |
| Python | 3.12.12 |
> 實測只用得到 cgroup 配的 16 個 vCPU。FLUX 模型極大(12B 參數),記憶體需求極高。
### 1.2 套件版本(皆為安裝當下可取得之最新組合)
| 套件 | 版本 |
| --- | --- |
| openvino | 2026.4.0 |
| nncf | 3.4.0 |
| optimum | 2.3.0 |
| optimum-intel | 2.2.0 |
| diffusers | 0.39.0 |
| transformers | 5.5.4 |
| tokenizers | 0.22.2 |
| huggingface-hub | 1.21.0 |
| torch | 2.14.1 (CPU) |
| pillow | 12.3.0 |
| psutil | 7.2.2 |
安裝指令:
```bash
pip install -U optimum optimum-intel openvino nncf \
diffusers transformers tokenizers huggingface-hub \
torch pillow psutil
```
> `optimum-intel 2.2.0` 對 `transformers` 宣告 `<5.6`、對 `huggingface-hub` 宣告 `<1.22`,因此在保持「各套件最新」的前提下,這是唯一能讓 optimum-cli / OVQuantizer / OVFluxPipeline 全部匯入成功的組合(diffusers 0.40.0 要求 `huggingface-hub>=1.23`,會與 optimum-intel 衝突,故取 diffusers 0.39.0)。
---
## 2. 轉換流程
### 2.1 步驟一:optimum-cli 匯出 FP16
```bash
optimum-cli export openvino \
-m black-forest-labs/FLUX.1-schnell \
--task text-to-image \
--library diffusers \
--weight-format fp16 \
/home/user/app/flux-schnell-ov-fp16
```
輸出(`flux-schnell-ov-fp16/`,共 **32175.7 MB ≈ 31.4 GB**):
```
model_index.json scheduler/ text_encoder/ text_encoder_2/
tokenizer/ tokenizer_2/ transformer/ vae_decoder/ vae_encoder/
```
`transformer/openvino_model.bin` 單獨就佔大多數空間(≈ 24 GB)。
### 2.2 步驟二:NNCF weight-only INT4(`quantize_int4_flux.py`)
```python
int4 = dict(bits=4, sym=False, group_size=128, group_size_fallback="adjust", ratio=1.0)
quantization_config = OVPipelineQuantizationConfig(
quantization_configs={
"transformer": OVWeightQuantizationConfig(**int4),
"text_encoder": OVWeightQuantizationConfig(**int4),
"text_encoder_2": OVWeightQuantizationConfig(**int4),
},
default_config=OVWeightQuantizationConfig(bits=8),
)
ov_config = OVConfig(quantization_config=quantization_config)
model = OVFluxPipeline.from_pretrained(fp16_dir, device="CPU")
OVQuantizer(model=model).quantize(ov_config=ov_config, save_directory=int4_dir)
```
- 量化耗時:**97.6 s**(不需要 calibration data,weight-only 為 data-free)
- 輸出 `flux-schnell-ov-int4/openvino_config.json` 記錄 `dtype: int4_int4`,transformer / text_encoder / text_encoder_2 皆為 `bits=4, sym=false, group_size=128, group_size_fallback=adjust, ratio=1.0`
### 2.3 模型大小比較
| Component | FP16 (MB) | INT4 (MB) | 壓縮 |
| --- | ---: | ---: | ---: |
| transformer (INT4) | ~24000 | ~5000 | ~4.8× |
| text_encoder (INT4) | ~240 | ~80 | ~3.0× |
| text_encoder_2 (INT4) | ~7500 | ~3000 | ~2.5× |
| vae_decoder (INT8) | ~100 | ~50 | ~2.0× |
| vae_encoder (INT8) | ~70 | ~35 | ~2.0× |
| **Pipeline 總計** | **32175.7** | **8511.4** | **3.78×**(-73.5%) |
整體磁碟用量 31.4 GB → 8.3 GB。
---
## 3. 推論實驗
### 3.1 設定
- `OVFluxPipeline.from_pretrained(int4_dir, compile=True, device="CPU")`
- 1024×1024,4 steps,guidance_scale = 0.0,max_sequence_length = 256,固定 seed 42–46
- 對照組:同 prompt / 同 seed,512×512
- 每一步以 diffusers `callback_on_step_end` 記錄 `perf_counter` 時間差與 `psutil` RSS
### 3.2 1024×1024 結果(4 steps, CFG 0.0)
| 影像 | seed | 總時間 | 單步平均 | 單步中位數 | 單步 min/max | 峰值 RSS |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| 01_hanfu | 42 | 76.69 s | 18.665 s | 16.055 s | 15.492 / 27.059 s | 34481 MB |
| 02_astronaut | 43 | 65.74 s | 16.014 s | 15.850 s | 15.000 / 17.355 s | 36888 MB |
| 03_taipei | 44 | 63.59 s | 15.471 s | 15.491 s | 15.211 / 15.693 s | 36892 MB |
| 04_shiba | 45 | 64.95 s | 15.764 s | 15.741 s | 15.572 / 16.004 s | 36897 MB |
| 05_ink | 46 | 63.83 s | 15.484 s | 15.405 s | 14.764 / 16.362 s | 36899 MB |
| **平均** | | **66.96 s** | **16.280 s** | **15.708 s** | | **~36.4 GB** |
`load + compile=True`:**11.82 s**,結束時 RSS 3423 MB。
> 01_hanfu 的第 1 步 27.06 s 是第一張圖的暖機/頁面缺失成本(`rss` 從 3.4 GB 漲到 34.5 GB),後續穩定在 15–16 s/step。
### 3.3 512×512 對照組(4 steps, CFG 0.0)
| 影像 | 總時間 | 單步平均 | 單步中位數 |
| --- | ---: | ---: | ---: |
| 01_hanfu_512 | 20.39 s | 4.947 s | 4.768 s |
| 02_astronaut_512 | 20.09 s | 4.921 s | 4.800 s |
| 03_taipei_512 | 19.84 s | 4.831 s | 4.700 s |
| 04_shiba_512 | 20.31 s | 4.973 s | 4.834 s |
| 05_ink_512 | 20.49 s | 5.001 s | 4.813 s |
解析度降 4 倍(1024²→512²,像素 1/4)→ 單步時間降為 **~1/3.3**(16.3 s → 4.9 s),符合 FLUX 在 CPU 上以 attention/activation 記憶體頻寬受限的特徵。
### 3.4 記憶體
- `load+compile` 完成後 RSS:**3423 MB**(約為 8511 MB 權重的 0.4×,含 OpenVINO 執行期 + 權重快取)
- 連續生成後 RSS 上限:**~36.9 GB**(第一張圖暖機後穩定;OpenVINO CPU 會快取 dequant / 顯存池不會隨影像釋放)
- 容器 cgroup 上限 104 GB,滿足需求;一般機器建議 **64 GB 以上** 記憶體才能舒適運行此 INT4 FLUX pipeline。
### 3.5 產出圖
1024×1024(主組,4 steps, CFG 0.0):
| 01_hanfu (seed 42) | 02_astronaut (seed 43) |
| --- | --- |
|  |  |
| 03_taipei (seed 44) | 04_shiba (seed 45) |
| --- | --- |
|  |  |
| 05_ink (seed 46) | |
| --- | --- |
|  | |
512×512 對照組:
| 01_hanfu_512 | 02_astronaut_512 | 03_taipei_512 |
| --- | --- | --- |
|  |  |  |
| 04_shiba_512 | 05_ink_512 | |
| --- | --- | --- |
|  |  | |
Prompt 全部列在 `prompts.txt`(FLUX 使用 guidance_scale=0.0,故無 negative prompt)。
---
## 4. 結論
- INT4 weight-only 讓 FLUX.1-schnell pipeline 從 31.4 GB 縮到 8.3 GB(**-73.5%**),其中 transformer 壓到 1/4.8、text_encoder 1/3.0、text_encoder_2 1/2.5。
- 純 CPU(16 vCPU Xeon 8559C)上 1024×1024 / 4 steps / CFG 0.0 出一張圖約 **64–77 s**(15–18 s/step),load+compile 12 s 內完成。
- 512×512 約 20 s/張,單步 4.8–5.0 s。
- 記憶體需求極高(峰值 ~37 GB),主要來自 12B 參數的 transformer,即使 INT4 量化後仍需大量記憶體儲存反量化權重與 activation。
- 品質方面,INT4 weight-only(group 128, asym)在 10 張測試圖中未見明顯劣化,人臉、霓虹中文字、留白水墨都能正常生成。
## 5. 檔案
| 檔案 | 說明 |
| --- | --- |
| `inference_int4_flux.py` | 單張推論(txt2img),`OVFluxPipeline.from_pretrained(..., compile=True)` |
| `quantize_int4_flux.py` | FP16 → NNCF weight-only INT4(transformer + text_encoder + text_encoder_2 INT4,其餘 INT8) |
| `generate5_flux.py` | 5 組固定 seed 生成 + callback 記錄每 step 時間與 RSS,輸出 benchmark.json / prompts.txt |
| `REPORT.md` | 本報告 |
| `examples/*.png` | 10 張範例圖(5×1024 + 5×512) |
| `outputs/benchmark.json` | 完整機器資訊 + 每張圖每 step 的時間/RSS |
| `outputs/prompts.txt` | 5 組 prompt / negative prompt / seed | |