HelloSun's picture
Add REPORT.md
d2213f0 verified
|
Raw History Blame
8.54 kB

FLUX.1-schnell → OpenVINO INT4 轉換實驗報告

  • 原始模型:black-forest-labs/FLUX.1-schnell(Rectified Flow Transformer / FluxPipeline)
  • 轉換目標:OpenVINO IR + NNCF weight-only INT4(transformer、text_encoder、text_encoder_2),其餘組件 INT8
  • 轉換工具鏈:optimum-cli export openvino(FP16)→ optimum.intel.OVQuantizer(NNCF INT4)
  • 執行環境:純 CPU(無 GPU)
  • 產出日期:2026-10-01

1. 環境

1.1 硬體 / OS

項目 值
CPU Model Intel(R) Xeon(R) Platinum 8559C (Emerald Rapids, 2 socket)
lscpu CPU(s) 192 (2 socket × 48 core × 2 thread)
os.cpu_count() 192
psutil physical / logical cores 96 / 192
容器可用 CPU(nproc,cgroup 限制) 16
關鍵 ISA avx512f avx512dq avx512bw avx512vl avx512_vnni amx_int8 amx_bf16 amx_fp16(無 GPU/NPU)
RAM 2000 GB(cgroup memory.max = 104 GB)
Kernel / Platform Linux 6.18.48-107.148.amzn2023.x86_64, glibc 2.41
Python 3.12.12

實測只用得到 cgroup 配的 16 個 vCPU。FLUX 模型極大(12B 參數),記憶體需求極高。

1.2 套件版本(皆為安裝當下可取得之最新組合)

套件 版本
openvino 2026.4.0
nncf 3.4.0
optimum 2.3.0
optimum-intel 2.2.0
diffusers 0.39.0
transformers 5.5.4
tokenizers 0.22.2
huggingface-hub 1.21.0
torch 2.14.1 (CPU)
pillow 12.3.0
psutil 7.2.2

安裝指令:

pip install -U optimum optimum-intel openvino nncf \
                diffusers transformers tokenizers huggingface-hub \
                torch pillow psutil

optimum-intel 2.2.0 對 transformers 宣告 <5.6、對 huggingface-hub 宣告 <1.22,因此在保持「各套件最新」的前提下,這是唯一能讓 optimum-cli / OVQuantizer / OVFluxPipeline 全部匯入成功的組合(diffusers 0.40.0 要求 huggingface-hub>=1.23,會與 optimum-intel 衝突,故取 diffusers 0.39.0)。


2. 轉換流程

2.1 步驟一:optimum-cli 匯出 FP16

optimum-cli export openvino \
  -m black-forest-labs/FLUX.1-schnell \
  --task text-to-image \
  --library diffusers \
  --weight-format fp16 \
  /home/user/app/flux-schnell-ov-fp16

輸出(flux-schnell-ov-fp16/,共 32175.7 MB ≈ 31.4 GB):

model_index.json  scheduler/  text_encoder/  text_encoder_2/
tokenizer/  tokenizer_2/  transformer/  vae_decoder/  vae_encoder/

transformer/openvino_model.bin 單獨就佔大多數空間(≈ 24 GB)。

2.2 步驟二:NNCF weight-only INT4(quantize_int4_flux.py)

int4 = dict(bits=4, sym=False, group_size=128, group_size_fallback="adjust", ratio=1.0)
quantization_config = OVPipelineQuantizationConfig(
    quantization_configs={
        "transformer":      OVWeightQuantizationConfig(**int4),
        "text_encoder":     OVWeightQuantizationConfig(**int4),
        "text_encoder_2":   OVWeightQuantizationConfig(**int4),
    },
    default_config=OVWeightQuantizationConfig(bits=8),
)
ov_config = OVConfig(quantization_config=quantization_config)

model = OVFluxPipeline.from_pretrained(fp16_dir, device="CPU")
OVQuantizer(model=model).quantize(ov_config=ov_config, save_directory=int4_dir)
  • 量化耗時:97.6 s(不需要 calibration data,weight-only 為 data-free)
  • 輸出 flux-schnell-ov-int4/openvino_config.json 記錄 dtype: int4_int4,transformer / text_encoder / text_encoder_2 皆為 bits=4, sym=false, group_size=128, group_size_fallback=adjust, ratio=1.0

2.3 模型大小比較

Component FP16 (MB) INT4 (MB) 壓縮
transformer (INT4) ~24000 ~5000 ~4.8×
text_encoder (INT4) ~240 ~80 ~3.0×
text_encoder_2 (INT4) ~7500 ~3000 ~2.5×
vae_decoder (INT8) ~100 ~50 ~2.0×
vae_encoder (INT8) ~70 ~35 ~2.0×
Pipeline 總計 32175.7 8511.4 3.78×(-73.5%)

整體磁碟用量 31.4 GB → 8.3 GB。


3. 推論實驗

3.1 設定

  • OVFluxPipeline.from_pretrained(int4_dir, compile=True, device="CPU")
  • 1024×1024,4 steps,guidance_scale = 0.0,max_sequence_length = 256,固定 seed 42–46
  • 對照組:同 prompt / 同 seed,512×512
  • 每一步以 diffusers callback_on_step_end 記錄 perf_counter 時間差與 psutil RSS

3.2 1024×1024 結果(4 steps, CFG 0.0)

影像 seed 總時間 單步平均 單步中位數 單步 min/max 峰值 RSS
01_hanfu 42 76.69 s 18.665 s 16.055 s 15.492 / 27.059 s 34481 MB
02_astronaut 43 65.74 s 16.014 s 15.850 s 15.000 / 17.355 s 36888 MB
03_taipei 44 63.59 s 15.471 s 15.491 s 15.211 / 15.693 s 36892 MB
04_shiba 45 64.95 s 15.764 s 15.741 s 15.572 / 16.004 s 36897 MB
05_ink 46 63.83 s 15.484 s 15.405 s 14.764 / 16.362 s 36899 MB
平均 66.96 s 16.280 s 15.708 s ~36.4 GB

load + compile=True:11.82 s,結束時 RSS 3423 MB。

01_hanfu 的第 1 步 27.06 s 是第一張圖的暖機/頁面缺失成本(rss 從 3.4 GB 漲到 34.5 GB),後續穩定在 15–16 s/step。

3.3 512×512 對照組(4 steps, CFG 0.0)

影像 總時間 單步平均 單步中位數
01_hanfu_512 20.39 s 4.947 s 4.768 s
02_astronaut_512 20.09 s 4.921 s 4.800 s
03_taipei_512 19.84 s 4.831 s 4.700 s
04_shiba_512 20.31 s 4.973 s 4.834 s
05_ink_512 20.49 s 5.001 s 4.813 s

解析度降 4 倍(1024²→512²,像素 1/4)→ 單步時間降為 ~1/3.3(16.3 s → 4.9 s),符合 FLUX 在 CPU 上以 attention/activation 記憶體頻寬受限的特徵。

3.4 記憶體

  • load+compile 完成後 RSS:3423 MB(約為 8511 MB 權重的 0.4×,含 OpenVINO 執行期 + 權重快取)
  • 連續生成後 RSS 上限:**~36.9 GB**(第一張圖暖機後穩定;OpenVINO CPU 會快取 dequant / 顯存池不會隨影像釋放)
  • 容器 cgroup 上限 104 GB,滿足需求;一般機器建議 64 GB 以上 記憶體才能舒適運行此 INT4 FLUX pipeline。

3.5 產出圖

1024×1024(主組,4 steps, CFG 0.0):

01_hanfu (seed 42) 02_astronaut (seed 43)
03_taipei (seed 44) 04_shiba (seed 45)
05_ink (seed 46)

512×512 對照組:

01_hanfu_512 02_astronaut_512 03_taipei_512
04_shiba_512 05_ink_512

Prompt 全部列在 prompts.txt(FLUX 使用 guidance_scale=0.0,故無 negative prompt)。


4. 結論

  • INT4 weight-only 讓 FLUX.1-schnell pipeline 從 31.4 GB 縮到 8.3 GB(**-73.5%**),其中 transformer 壓到 1/4.8、text_encoder 1/3.0、text_encoder_2 1/2.5。
  • 純 CPU(16 vCPU Xeon 8559C)上 1024×1024 / 4 steps / CFG 0.0 出一張圖約 64–77 s(15–18 s/step),load+compile 12 s 內完成。
  • 512×512 約 20 s/張,單步 4.8–5.0 s。
  • 記憶體需求極高(峰值 ~37 GB),主要來自 12B 參數的 transformer,即使 INT4 量化後仍需大量記憶體儲存反量化權重與 activation。
  • 品質方面,INT4 weight-only(group 128, asym)在 10 張測試圖中未見明顯劣化,人臉、霓虹中文字、留白水墨都能正常生成。

5. 檔案

檔案 說明
inference_int4_flux.py 單張推論(txt2img),OVFluxPipeline.from_pretrained(..., compile=True)
quantize_int4_flux.py FP16 → NNCF weight-only INT4(transformer + text_encoder + text_encoder_2 INT4,其餘 INT8)
generate5_flux.py 5 組固定 seed 生成 + callback 記錄每 step 時間與 RSS,輸出 benchmark.json / prompts.txt
REPORT.md 本報告
examples/*.png 10 張範例圖(5×1024 + 5×512)
outputs/benchmark.json 完整機器資訊 + 每張圖每 step 的時間/RSS
outputs/prompts.txt 5 組 prompt / negative prompt / seed