FLUX.1-schnell · OpenVINO INT4 (CPU)

black-forest-labs/FLUX.1-schnell 轉成 OpenVINO IR, transformer、text_encoder、text_encoder_2 以 NNCF 做 weight-only INT4(asymmetric, group_size 128), 其餘組件 INT8。不需要 GPU,純 CPU 即可生圖。

  • 權重體積 31.4 GB → 8.3 GB(-73.5%)
  • 1024×1024 / 4 steps / CFG 0.0:**~67 s 出一張(16.3 s/step)**(Xeon Platinum 8559C, 16 vCPU)
  • 512×512 / 4 steps:**~20 s 出一張(4.9 s/step)**
  • load + compile:11.8 s

轉換環境:openvino 2026.4.0 / nncf 3.4.0 / optimum 2.3.0 / optimum-intel 2.2.0 / diffusers 0.39.0 / transformers 5.5.4 / tokenizers 0.22.2 / huggingface-hub 1.21.0 / torch 2.14.1 / pillow 12.3.0 / psutil 7.2.2


模型頁展示(10 張)

1024 × 1024(4 steps, CFG 0.0, max_sequence_length=256)

seed 42 · 01_hanfu seed 43 · 02_astronaut
seed 44 · 03_taipei seed 45 · 04_shiba
seed 46 · 05_ink

512 × 512 對照組(同 prompt / 同 seed)

01_hanfu_512 02_astronaut_512 03_taipei_512
04_shiba_512 05_ink_512

Prompts 見 outputs/prompts.txt。


安裝

pip install -U optimum optimum-intel openvino nncf \
                diffusers transformers tokenizers huggingface-hub \
                torch pillow psutil

用法

import torch
from optimum.intel import OVFluxPipeline

pipe = OVFluxPipeline.from_pretrained(
    "HelloSun/FLUX.1-schnell-OpenVINO-INT4",
    compile=True,
    device="CPU",
)

prompt = (
    "Young Chinese woman in red Hanfu, intricate embroidery, impeccable makeup, "
    "red floral forehead pattern, elaborate high bun, golden phoenix headdress, "
    "soft-lit outdoor night background, silhouetted tiered pagoda, blurred colorful "
    "distant lights, photorealistic, ultra detailed, 8k"
)
generator = torch.Generator(device="cpu").manual_seed(42)

image = pipe(
    prompt=prompt,
    negative_prompt="",
    width=1024, height=1024,
    num_inference_steps=4,
    guidance_scale=0.0,
    max_sequence_length=256,
    generator=generator,
).images[0]
image.save("hanfu.png")

完整腳本:

檔案 用途
inference_int4_flux.py 單張 txt2img 推論
quantize_int4_flux.py 由 FP16 pipeline 產生 INT4 pipeline(OVQuantizer + NNCF)
generate5_flux.py 5 組固定 seed 批次生成,callback 記錄每 step 時間與 RSS
REPORT.md 完整轉換與實驗報告
outputs/benchmark.json 機器資訊 + 每 step 時間/RSS 原始數據
outputs/prompts.txt 全部 prompt / negative prompt / seed

轉換步驟(重現本 repo)

# 1) 匯出 FP16
optimum-cli export openvino -m black-forest-labs/FLUX.1-schnell \
  --task text-to-image --library diffusers --weight-format fp16 \
  ./flux-schnell-ov-fp16

# 2) NNCF weight-only INT4(transformer + text_encoder + text_encoder_2 INT4,其餘 INT8)
python quantize_int4_flux.py \
  --model_path ./flux-schnell-ov-fp16 \
  --output_path ./flux-schnell-ov-int4

量化設定(quantize_int4_flux.py 內):

int4 = dict(bits=4, sym=False, group_size=128, group_size_fallback="adjust", ratio=1.0)
OVPipelineQuantizationConfig(
    quantization_configs={
        "transformer":      OVWeightQuantizationConfig(**int4),
        "text_encoder":     OVWeightQuantizationConfig(**int4),
        "text_encoder_2":   OVWeightQuantizationConfig(**int4),
    },
    default_config=OVWeightQuantizationConfig(bits=8),
)

實驗數據

環境:Intel Xeon Platinum 8559C(2 socket / 96 core / 192 thread,容器 cgroup 限 16 vCPU)、 AVX-512、2 TB RAM、無 GPU。lscpu CPU(s)=192,os.cpu_count()=192,OpenVINO 2026.4.0。

項目 數值
load + compile(compile=True) 11.82 s(結束時 RSS 3423 MB)
1024×1024 / 4 steps 單張總時間 63.59 – 76.69 s(平均 66.96 s)
1024×1024 單步時間 平均 16.28 s、中位數 15.71 s、範圍 14.76 – 27.06 s(暖機除外)
512×512 / 4 steps 單張總時間 19.84 – 20.49 s(平均 20.22 s)
512×512 單步時間 平均 4.93 s、中位數 4.78 s
峰值 RSS(連續生成) ~36.9 GB
FP16 pipeline 大小 32175.7 MB
INT4 pipeline 大小 8511.4 MB(-73.5%)
量化時間(weight-only, data-free) 97.6 s

逐張明細:

影像 seed 解析度 steps CFG 總時間 單步平均 峰值 RSS
01_hanfu 42 1024² 4 0.0 76.69 s 18.665 s 34481 MB
02_astronaut 43 1024² 4 0.0 65.74 s 16.014 s 36888 MB
03_taipei 44 1024² 4 0.0 63.59 s 15.471 s 36892 MB
04_shiba 45 1024² 4 0.0 64.95 s 15.764 s 36897 MB
05_ink 46 1024² 4 0.0 63.83 s 15.484 s 36899 MB
01_hanfu_512 42 512² 4 0.0 20.39 s 4.947 s 36903 MB
02_astronaut_512 43 512² 4 0.0 20.09 s 4.921 s 36912 MB
03_taipei_512 44 512² 4 0.0 19.84 s 4.831 s 36912 MB
04_shiba_512 45 512² 4 0.0 20.31 s 4.973 s 36912 MB
05_ink_512 46 512² 4 0.0 20.49 s 5.001 s 36912 MB

建議參數

沿用上游模型卡建議:steps 1–4、CFG 0.0(FLUX 為 rectified flow,不使用 CFG)、max_sequence_length=256。 本 repo 範例採 4 steps。

Credits

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HelloSun/FLUX.1-schnell-OpenVINO-INT4

Finetuned
(73)
this model