Download REPORT.md from HelloSun/FLUX.1-schnell-OpenVINO-INT4: direct link, hf CLI and curl.
- Browser
- Download file 8.54 kB
-
https://huggingface.co/HelloSun/FLUX.1-schnell-OpenVINO-INT4/resolve/ff258978ce7ebb2c05573c45827595c81807dca7/REPORT.md
- Command line
-
hf download hf://HelloSun/FLUX.1-schnell-OpenVINO-INT4@ff258978ce7ebb2c05573c45827595c81807dca7/REPORT.md
-
curl -L -o REPORT.md https://huggingface.co/HelloSun/FLUX.1-schnell-OpenVINO-INT4/resolve/ff258978ce7ebb2c05573c45827595c81807dca7/REPORT.md
FLUX.1-schnell → OpenVINO INT4 轉換實驗報告
- 原始模型:black-forest-labs/FLUX.1-schnell(Rectified Flow Transformer / FluxPipeline)
- 轉換目標:OpenVINO IR + NNCF weight-only INT4(transformer、text_encoder、text_encoder_2),其餘組件 INT8
- 轉換工具鏈:
optimum-cli export openvino(FP16)→optimum.intel.OVQuantizer(NNCF INT4) - 執行環境:純 CPU(無 GPU)
- 產出日期:2026-10-01
1. 環境
1.1 硬體 / OS
| 項目 | 值 |
|---|---|
| CPU Model | Intel(R) Xeon(R) Platinum 8559C (Emerald Rapids, 2 socket) |
lscpu CPU(s) |
192 (2 socket × 48 core × 2 thread) |
os.cpu_count() |
192 |
| psutil physical / logical cores | 96 / 192 |
容器可用 CPU(nproc,cgroup 限制) |
16 |
| 關鍵 ISA | avx512f avx512dq avx512bw avx512vl avx512_vnni amx_int8 amx_bf16 amx_fp16(無 GPU/NPU) |
| RAM | 2000 GB(cgroup memory.max = 104 GB) |
| Kernel / Platform | Linux 6.18.48-107.148.amzn2023.x86_64, glibc 2.41 |
| Python | 3.12.12 |
實測只用得到 cgroup 配的 16 個 vCPU。FLUX 模型極大(12B 參數),記憶體需求極高。
1.2 套件版本(皆為安裝當下可取得之最新組合)
| 套件 | 版本 |
|---|---|
| openvino | 2026.4.0 |
| nncf | 3.4.0 |
| optimum | 2.3.0 |
| optimum-intel | 2.2.0 |
| diffusers | 0.39.0 |
| transformers | 5.5.4 |
| tokenizers | 0.22.2 |
| huggingface-hub | 1.21.0 |
| torch | 2.14.1 (CPU) |
| pillow | 12.3.0 |
| psutil | 7.2.2 |
安裝指令:
pip install -U optimum optimum-intel openvino nncf \
diffusers transformers tokenizers huggingface-hub \
torch pillow psutil
optimum-intel 2.2.0對transformers宣告<5.6、對huggingface-hub宣告<1.22,因此在保持「各套件最新」的前提下,這是唯一能讓 optimum-cli / OVQuantizer / OVFluxPipeline 全部匯入成功的組合(diffusers 0.40.0 要求huggingface-hub>=1.23,會與 optimum-intel 衝突,故取 diffusers 0.39.0)。
2. 轉換流程
2.1 步驟一:optimum-cli 匯出 FP16
optimum-cli export openvino \
-m black-forest-labs/FLUX.1-schnell \
--task text-to-image \
--library diffusers \
--weight-format fp16 \
/home/user/app/flux-schnell-ov-fp16
輸出(flux-schnell-ov-fp16/,共 32175.7 MB ≈ 31.4 GB):
model_index.json scheduler/ text_encoder/ text_encoder_2/
tokenizer/ tokenizer_2/ transformer/ vae_decoder/ vae_encoder/
transformer/openvino_model.bin 單獨就佔大多數空間(≈ 24 GB)。
2.2 步驟二:NNCF weight-only INT4(quantize_int4_flux.py)
int4 = dict(bits=4, sym=False, group_size=128, group_size_fallback="adjust", ratio=1.0)
quantization_config = OVPipelineQuantizationConfig(
quantization_configs={
"transformer": OVWeightQuantizationConfig(**int4),
"text_encoder": OVWeightQuantizationConfig(**int4),
"text_encoder_2": OVWeightQuantizationConfig(**int4),
},
default_config=OVWeightQuantizationConfig(bits=8),
)
ov_config = OVConfig(quantization_config=quantization_config)
model = OVFluxPipeline.from_pretrained(fp16_dir, device="CPU")
OVQuantizer(model=model).quantize(ov_config=ov_config, save_directory=int4_dir)
- 量化耗時:97.6 s(不需要 calibration data,weight-only 為 data-free)
- 輸出
flux-schnell-ov-int4/openvino_config.json記錄dtype: int4_int4,transformer / text_encoder / text_encoder_2 皆為bits=4, sym=false, group_size=128, group_size_fallback=adjust, ratio=1.0
2.3 模型大小比較
| Component | FP16 (MB) | INT4 (MB) | 壓縮 |
|---|---|---|---|
| transformer (INT4) | ~24000 | ~5000 | ~4.8× |
| text_encoder (INT4) | ~240 | ~80 | ~3.0× |
| text_encoder_2 (INT4) | ~7500 | ~3000 | ~2.5× |
| vae_decoder (INT8) | ~100 | ~50 | ~2.0× |
| vae_encoder (INT8) | ~70 | ~35 | ~2.0× |
| Pipeline 總計 | 32175.7 | 8511.4 | 3.78×(-73.5%) |
整體磁碟用量 31.4 GB → 8.3 GB。
3. 推論實驗
3.1 設定
OVFluxPipeline.from_pretrained(int4_dir, compile=True, device="CPU")- 1024×1024,4 steps,guidance_scale = 0.0,max_sequence_length = 256,固定 seed 42–46
- 對照組:同 prompt / 同 seed,512×512
- 每一步以 diffusers
callback_on_step_end記錄perf_counter時間差與psutilRSS
3.2 1024×1024 結果(4 steps, CFG 0.0)
| 影像 | seed | 總時間 | 單步平均 | 單步中位數 | 單步 min/max | 峰值 RSS |
|---|---|---|---|---|---|---|
| 01_hanfu | 42 | 76.69 s | 18.665 s | 16.055 s | 15.492 / 27.059 s | 34481 MB |
| 02_astronaut | 43 | 65.74 s | 16.014 s | 15.850 s | 15.000 / 17.355 s | 36888 MB |
| 03_taipei | 44 | 63.59 s | 15.471 s | 15.491 s | 15.211 / 15.693 s | 36892 MB |
| 04_shiba | 45 | 64.95 s | 15.764 s | 15.741 s | 15.572 / 16.004 s | 36897 MB |
| 05_ink | 46 | 63.83 s | 15.484 s | 15.405 s | 14.764 / 16.362 s | 36899 MB |
| 平均 | 66.96 s | 16.280 s | 15.708 s | ~36.4 GB |
load + compile=True:11.82 s,結束時 RSS 3423 MB。
01_hanfu 的第 1 步 27.06 s 是第一張圖的暖機/頁面缺失成本(
rss從 3.4 GB 漲到 34.5 GB),後續穩定在 15–16 s/step。
3.3 512×512 對照組(4 steps, CFG 0.0)
| 影像 | 總時間 | 單步平均 | 單步中位數 |
|---|---|---|---|
| 01_hanfu_512 | 20.39 s | 4.947 s | 4.768 s |
| 02_astronaut_512 | 20.09 s | 4.921 s | 4.800 s |
| 03_taipei_512 | 19.84 s | 4.831 s | 4.700 s |
| 04_shiba_512 | 20.31 s | 4.973 s | 4.834 s |
| 05_ink_512 | 20.49 s | 5.001 s | 4.813 s |
解析度降 4 倍(1024²→512²,像素 1/4)→ 單步時間降為 ~1/3.3(16.3 s → 4.9 s),符合 FLUX 在 CPU 上以 attention/activation 記憶體頻寬受限的特徵。
3.4 記憶體
load+compile完成後 RSS:3423 MB(約為 8511 MB 權重的 0.4×,含 OpenVINO 執行期 + 權重快取)- 連續生成後 RSS 上限:**~36.9 GB**(第一張圖暖機後穩定;OpenVINO CPU 會快取 dequant / 顯存池不會隨影像釋放)
- 容器 cgroup 上限 104 GB,滿足需求;一般機器建議 64 GB 以上 記憶體才能舒適運行此 INT4 FLUX pipeline。
3.5 產出圖
1024×1024(主組,4 steps, CFG 0.0):
512×512 對照組:
Prompt 全部列在 prompts.txt(FLUX 使用 guidance_scale=0.0,故無 negative prompt)。
4. 結論
- INT4 weight-only 讓 FLUX.1-schnell pipeline 從 31.4 GB 縮到 8.3 GB(**-73.5%**),其中 transformer 壓到 1/4.8、text_encoder 1/3.0、text_encoder_2 1/2.5。
- 純 CPU(16 vCPU Xeon 8559C)上 1024×1024 / 4 steps / CFG 0.0 出一張圖約 64–77 s(15–18 s/step),load+compile 12 s 內完成。
- 512×512 約 20 s/張,單步 4.8–5.0 s。
- 記憶體需求極高(峰值 ~37 GB),主要來自 12B 參數的 transformer,即使 INT4 量化後仍需大量記憶體儲存反量化權重與 activation。
- 品質方面,INT4 weight-only(group 128, asym)在 10 張測試圖中未見明顯劣化,人臉、霓虹中文字、留白水墨都能正常生成。
5. 檔案
| 檔案 | 說明 |
|---|---|
inference_int4_flux.py |
單張推論(txt2img),OVFluxPipeline.from_pretrained(..., compile=True) |
quantize_int4_flux.py |
FP16 → NNCF weight-only INT4(transformer + text_encoder + text_encoder_2 INT4,其餘 INT8) |
generate5_flux.py |
5 組固定 seed 生成 + callback 記錄每 step 時間與 RSS,輸出 benchmark.json / prompts.txt |
REPORT.md |
本報告 |
examples/*.png |
10 張範例圖(5×1024 + 5×512) |
outputs/benchmark.json |
完整機器資訊 + 每張圖每 step 的時間/RSS |
outputs/prompts.txt |
5 組 prompt / negative prompt / seed |









