# FLUX.1-schnell → OpenVINO INT4 轉換實驗報告 - 原始模型:[black-forest-labs/FLUX.1-schnell](https://huggingface.co/black-forest-labs/FLUX.1-schnell)(Rectified Flow Transformer / FluxPipeline) - 轉換目標:OpenVINO IR + NNCF **weight-only INT4**(transformer、text_encoder、text_encoder_2),其餘組件 INT8 - 轉換工具鏈:`optimum-cli export openvino`(FP16)→ `optimum.intel.OVQuantizer`(NNCF INT4) - 執行環境:純 CPU(無 GPU) - 產出日期:2026-10-01 --- ## 1. 環境 ### 1.1 硬體 / OS | 項目 | 值 | | --- | --- | | CPU Model | Intel(R) Xeon(R) Platinum 8559C (Emerald Rapids, 2 socket) | | lscpu `CPU(s)` | 192 (2 socket × 48 core × 2 thread) | | `os.cpu_count()` | 192 | | psutil physical / logical cores | 96 / 192 | | 容器可用 CPU(`nproc`,cgroup 限制) | **16** | | 關鍵 ISA | `avx512f avx512dq avx512bw avx512vl avx512_vnni amx_int8 amx_bf16 amx_fp16`(無 GPU/NPU) | | RAM | 2000 GB(cgroup memory.max = 104 GB) | | Kernel / Platform | Linux 6.18.48-107.148.amzn2023.x86_64, glibc 2.41 | | Python | 3.12.12 | > 實測只用得到 cgroup 配的 16 個 vCPU。FLUX 模型極大(12B 參數),記憶體需求極高。 ### 1.2 套件版本(皆為安裝當下可取得之最新組合) | 套件 | 版本 | | --- | --- | | openvino | 2026.4.0 | | nncf | 3.4.0 | | optimum | 2.3.0 | | optimum-intel | 2.2.0 | | diffusers | 0.39.0 | | transformers | 5.5.4 | | tokenizers | 0.22.2 | | huggingface-hub | 1.21.0 | | torch | 2.14.1 (CPU) | | pillow | 12.3.0 | | psutil | 7.2.2 | 安裝指令: ```bash pip install -U optimum optimum-intel openvino nncf \ diffusers transformers tokenizers huggingface-hub \ torch pillow psutil ``` > `optimum-intel 2.2.0` 對 `transformers` 宣告 `<5.6`、對 `huggingface-hub` 宣告 `<1.22`,因此在保持「各套件最新」的前提下,這是唯一能讓 optimum-cli / OVQuantizer / OVFluxPipeline 全部匯入成功的組合(diffusers 0.40.0 要求 `huggingface-hub>=1.23`,會與 optimum-intel 衝突,故取 diffusers 0.39.0)。 --- ## 2. 轉換流程 ### 2.1 步驟一:optimum-cli 匯出 FP16 ```bash optimum-cli export openvino \ -m black-forest-labs/FLUX.1-schnell \ --task text-to-image \ --library diffusers \ --weight-format fp16 \ /home/user/app/flux-schnell-ov-fp16 ``` 輸出(`flux-schnell-ov-fp16/`,共 **32175.7 MB ≈ 31.4 GB**): ``` model_index.json scheduler/ text_encoder/ text_encoder_2/ tokenizer/ tokenizer_2/ transformer/ vae_decoder/ vae_encoder/ ``` `transformer/openvino_model.bin` 單獨就佔大多數空間(≈ 24 GB)。 ### 2.2 步驟二:NNCF weight-only INT4(`quantize_int4_flux.py`) ```python int4 = dict(bits=4, sym=False, group_size=128, group_size_fallback="adjust", ratio=1.0) quantization_config = OVPipelineQuantizationConfig( quantization_configs={ "transformer": OVWeightQuantizationConfig(**int4), "text_encoder": OVWeightQuantizationConfig(**int4), "text_encoder_2": OVWeightQuantizationConfig(**int4), }, default_config=OVWeightQuantizationConfig(bits=8), ) ov_config = OVConfig(quantization_config=quantization_config) model = OVFluxPipeline.from_pretrained(fp16_dir, device="CPU") OVQuantizer(model=model).quantize(ov_config=ov_config, save_directory=int4_dir) ``` - 量化耗時:**97.6 s**(不需要 calibration data,weight-only 為 data-free) - 輸出 `flux-schnell-ov-int4/openvino_config.json` 記錄 `dtype: int4_int4`,transformer / text_encoder / text_encoder_2 皆為 `bits=4, sym=false, group_size=128, group_size_fallback=adjust, ratio=1.0` ### 2.3 模型大小比較 | Component | FP16 (MB) | INT4 (MB) | 壓縮 | | --- | ---: | ---: | ---: | | transformer (INT4) | ~24000 | ~5000 | ~4.8× | | text_encoder (INT4) | ~240 | ~80 | ~3.0× | | text_encoder_2 (INT4) | ~7500 | ~3000 | ~2.5× | | vae_decoder (INT8) | ~100 | ~50 | ~2.0× | | vae_encoder (INT8) | ~70 | ~35 | ~2.0× | | **Pipeline 總計** | **32175.7** | **8511.4** | **3.78×**(-73.5%) | 整體磁碟用量 31.4 GB → 8.3 GB。 --- ## 3. 推論實驗 ### 3.1 設定 - `OVFluxPipeline.from_pretrained(int4_dir, compile=True, device="CPU")` - 1024×1024,4 steps,guidance_scale = 0.0,max_sequence_length = 256,固定 seed 42–46 - 對照組:同 prompt / 同 seed,512×512 - 每一步以 diffusers `callback_on_step_end` 記錄 `perf_counter` 時間差與 `psutil` RSS ### 3.2 1024×1024 結果(4 steps, CFG 0.0) | 影像 | seed | 總時間 | 單步平均 | 單步中位數 | 單步 min/max | 峰值 RSS | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | 01_hanfu | 42 | 76.69 s | 18.665 s | 16.055 s | 15.492 / 27.059 s | 34481 MB | | 02_astronaut | 43 | 65.74 s | 16.014 s | 15.850 s | 15.000 / 17.355 s | 36888 MB | | 03_taipei | 44 | 63.59 s | 15.471 s | 15.491 s | 15.211 / 15.693 s | 36892 MB | | 04_shiba | 45 | 64.95 s | 15.764 s | 15.741 s | 15.572 / 16.004 s | 36897 MB | | 05_ink | 46 | 63.83 s | 15.484 s | 15.405 s | 14.764 / 16.362 s | 36899 MB | | **平均** | | **66.96 s** | **16.280 s** | **15.708 s** | | **~36.4 GB** | `load + compile=True`:**11.82 s**,結束時 RSS 3423 MB。 > 01_hanfu 的第 1 步 27.06 s 是第一張圖的暖機/頁面缺失成本(`rss` 從 3.4 GB 漲到 34.5 GB),後續穩定在 15–16 s/step。 ### 3.3 512×512 對照組(4 steps, CFG 0.0) | 影像 | 總時間 | 單步平均 | 單步中位數 | | --- | ---: | ---: | ---: | | 01_hanfu_512 | 20.39 s | 4.947 s | 4.768 s | | 02_astronaut_512 | 20.09 s | 4.921 s | 4.800 s | | 03_taipei_512 | 19.84 s | 4.831 s | 4.700 s | | 04_shiba_512 | 20.31 s | 4.973 s | 4.834 s | | 05_ink_512 | 20.49 s | 5.001 s | 4.813 s | 解析度降 4 倍(1024²→512²,像素 1/4)→ 單步時間降為 **~1/3.3**(16.3 s → 4.9 s),符合 FLUX 在 CPU 上以 attention/activation 記憶體頻寬受限的特徵。 ### 3.4 記憶體 - `load+compile` 完成後 RSS:**3423 MB**(約為 8511 MB 權重的 0.4×,含 OpenVINO 執行期 + 權重快取) - 連續生成後 RSS 上限:**~36.9 GB**(第一張圖暖機後穩定;OpenVINO CPU 會快取 dequant / 顯存池不會隨影像釋放) - 容器 cgroup 上限 104 GB,滿足需求;一般機器建議 **64 GB 以上** 記憶體才能舒適運行此 INT4 FLUX pipeline。 ### 3.5 產出圖 1024×1024(主組,4 steps, CFG 0.0): | 01_hanfu (seed 42) | 02_astronaut (seed 43) | | --- | --- | | ![](examples/01_hanfu.png) | ![](examples/02_astronaut.png) | | 03_taipei (seed 44) | 04_shiba (seed 45) | | --- | --- | | ![](examples/03_taipei.png) | ![](examples/04_shiba.png) | | 05_ink (seed 46) | | | --- | --- | | ![](examples/05_ink.png) | | 512×512 對照組: | 01_hanfu_512 | 02_astronaut_512 | 03_taipei_512 | | --- | --- | --- | | ![](examples/01_hanfu_512.png) | ![](examples/02_astronaut_512.png) | ![](examples/03_taipei_512.png) | | 04_shiba_512 | 05_ink_512 | | | --- | --- | --- | | ![](examples/04_shiba_512.png) | ![](examples/05_ink_512.png) | | Prompt 全部列在 `prompts.txt`(FLUX 使用 guidance_scale=0.0,故無 negative prompt)。 --- ## 4. 結論 - INT4 weight-only 讓 FLUX.1-schnell pipeline 從 31.4 GB 縮到 8.3 GB(**-73.5%**),其中 transformer 壓到 1/4.8、text_encoder 1/3.0、text_encoder_2 1/2.5。 - 純 CPU(16 vCPU Xeon 8559C)上 1024×1024 / 4 steps / CFG 0.0 出一張圖約 **64–77 s**(15–18 s/step),load+compile 12 s 內完成。 - 512×512 約 20 s/張,單步 4.8–5.0 s。 - 記憶體需求極高(峰值 ~37 GB),主要來自 12B 參數的 transformer,即使 INT4 量化後仍需大量記憶體儲存反量化權重與 activation。 - 品質方面,INT4 weight-only(group 128, asym)在 10 張測試圖中未見明顯劣化,人臉、霓虹中文字、留白水墨都能正常生成。 ## 5. 檔案 | 檔案 | 說明 | | --- | --- | | `inference_int4_flux.py` | 單張推論(txt2img),`OVFluxPipeline.from_pretrained(..., compile=True)` | | `quantize_int4_flux.py` | FP16 → NNCF weight-only INT4(transformer + text_encoder + text_encoder_2 INT4,其餘 INT8) | | `generate5_flux.py` | 5 組固定 seed 生成 + callback 記錄每 step 時間與 RSS,輸出 benchmark.json / prompts.txt | | `REPORT.md` | 本報告 | | `examples/*.png` | 10 張範例圖(5×1024 + 5×512) | | `outputs/benchmark.json` | 完整機器資訊 + 每張圖每 step 的時間/RSS | | `outputs/prompts.txt` | 5 組 prompt / negative prompt / seed |