Qwen3.5-27B-NVFP4-GPTQ

GPTQ (Hessian-corrected PTQ) NVFP4 quantization of Qwen/Qwen3.5-27B in vLLM compressed-tensors nvfp4-pack-quantized format, with calibrated input global scales (W4A4-servable, default load, SM100+). This is the strong PTQ baseline of the stage-3 W4A4-training study — same lattice and serving path as weili-0234/Qwen3.5-27B-NVFP4-RTN, but with GPTQ weight error correction and activation-scale calibration on a fixed calibration set.

How this checkpoint was produced — exact reproduction

Tooling llm-compressor 0.12.0 (GPTQModifier), torch 2.10.0+cu128, transformers 5.9.0, Python 3.12.3
Base model Qwen/Qwen3.5-27B (dense qwen3_5, bf16)
Recipe GPTQModifier(targets="Linear", scheme="NVFP4", ignore=["lm_head", "re:.*visual.*", "re:.*mtp.*", "re:.*embed_tokens.*"]) — identical module scope to the RTN and QAD checkpoints in this series (all attention + MLP linears; GDN conv1d untouched — GPTQModifier matches nn.Linear only). The NVFP4 scheme also calibrates static per-tensor input global scales from the same calibration pass.
Calibration Fixed 512-conversation set calib_512.jsonl (md5 d1ff8cce1d785f8e51b71eb237fb7a71): random.Random(42).sample(range(n_rows), 512), sorted, from the study's OpenPerfectBlend-derived ChatML train corpus (train jsonl md5 0406bb3a7a482352360716a1bc5e9e04); rendered with the model's stock chat template, add_generation_prompt=False; max_seq_length=8192, num_calibration_samples=512
Script experiments/w4a4-qwen35-27b/jbom/gptq_27b.py in the study workspace (same recipe as weili-0234/Qwen3.5-9B-MXFP4-GPTQ, scheme swapped to NVFP4)
Hardware 1×B200 (jbom-03), ~2.2 h wall
Role in the study Strong-PTQ baseline row of the stage-3 Qwen3.5-27B 2×2 (W4A4-training vs W4A16-training × MXFP4 vs NVFP4). Companions: Qwen3.5-27B-NVFP4-RTN (weak PTQ) and the NVFP4 QAD arms as they complete.
# gptq_27b.py NVFP4  (llmcompressor==0.12.0, transformers==5.9.0)
from datasets import load_dataset
from transformers import AutoTokenizer, Qwen3_5ForConditionalGeneration
from llmcompressor import oneshot
from llmcompressor.modifiers.gptq import GPTQModifier

model = Qwen3_5ForConditionalGeneration.from_pretrained("Qwen/Qwen3.5-27B", dtype="bfloat16")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3.5-27B")
ds = load_dataset("json", data_files="calib_512.jsonl", split="train")
recipe = GPTQModifier(
    targets="Linear",
    scheme="NVFP4",
    ignore=["lm_head", "re:.*visual.*", "re:.*mtp.*", "re:.*embed_tokens.*"],
)
oneshot(model=model, tokenizer=tokenizer, dataset=ds, recipe=recipe,
        max_seq_length=8192, num_calibration_samples=512)
model.save_pretrained("Qwen3.5-27B-NVFP4-GPTQ", save_compressed=True)
tokenizer.save_pretrained("Qwen3.5-27B-NVFP4-GPTQ")

Serving

vllm serve weili-0234/Qwen3.5-27B-NVFP4-GPTQ    # W4A4 (native NVFP4 activations, SM100+)

Serving-KL vs BF16 (2026-07-24)

Mean top-20 token KL vs the BF16 teacher (Qwen/Qwen3.5-27B) over a held-out 256-conversation corpus (818,944 scored tokens), vLLM on 1×B200, W4A4 serving.

checkpoint KL vs BF16 top-1 agree
NVFP4-RTN (QATFactory export, running-max input scales) 0.0375 —
this repo (NVFP4-GPTQ, llm-compressor calibrated input scales) 0.0982 0.917

Caveat — this artifact does NOT beat RTN on serving-KL for NVFP4. Unlike MXFP4 (where GPTQ recovers ~23% of the RTN gap), the llm-compressor NVFP4 pipeline lands 2.6× worse than the QATFactory RTN export at W4A4 serving. The leading hypothesis is the input-activation global-scale calibration (static observer scales here vs running-max scales in the RTN export) dominating NVFP4's small weight-quantization error; the GPTQ weight update itself may be neutral-to-positive. Treat this checkpoint as a faithful record of the standard llm-compressor GPTQ-NVFP4 recipe at 27B rather than as a strengthened PTQ baseline.

A 7-benchmark evaluation (RULER, AIME'25 avg@4, GPQA-Diamond, MMLU-Pro, HumanEval, IFEval, Pile-10k/WikiText perplexity) against the RTN and QAD checkpoints is in progress and will be added to this card.

Downloads last month
21
Safetensors
Model size
17B params
Tensor type
F32
·
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for weili-0234/Qwen3.5-27B-NVFP4-GPTQ

Base model

Qwen/Qwen3.5-27B
Quantized
(229)
this model