Qwen3.5-27B-NVFP4-GPTQ
GPTQ (Hessian-corrected PTQ) NVFP4 quantization of Qwen/Qwen3.5-27B in vLLM
compressed-tensors nvfp4-pack-quantized format, with calibrated input
global scales (W4A4-servable, default load, SM100+). This is the strong
PTQ baseline of the stage-3 W4A4-training study — same lattice and serving
path as weili-0234/Qwen3.5-27B-NVFP4-RTN, but with GPTQ weight error
correction and activation-scale calibration on a fixed calibration set.
How this checkpoint was produced — exact reproduction
| Tooling | llm-compressor 0.12.0 (GPTQModifier), torch 2.10.0+cu128, transformers 5.9.0, Python 3.12.3 |
| Base model | Qwen/Qwen3.5-27B (dense qwen3_5, bf16) |
| Recipe | GPTQModifier(targets="Linear", scheme="NVFP4", ignore=["lm_head", "re:.*visual.*", "re:.*mtp.*", "re:.*embed_tokens.*"]) — identical module scope to the RTN and QAD checkpoints in this series (all attention + MLP linears; GDN conv1d untouched — GPTQModifier matches nn.Linear only). The NVFP4 scheme also calibrates static per-tensor input global scales from the same calibration pass. |
| Calibration | Fixed 512-conversation set calib_512.jsonl (md5 d1ff8cce1d785f8e51b71eb237fb7a71): random.Random(42).sample(range(n_rows), 512), sorted, from the study's OpenPerfectBlend-derived ChatML train corpus (train jsonl md5 0406bb3a7a482352360716a1bc5e9e04); rendered with the model's stock chat template, add_generation_prompt=False; max_seq_length=8192, num_calibration_samples=512 |
| Script | experiments/w4a4-qwen35-27b/jbom/gptq_27b.py in the study workspace (same recipe as weili-0234/Qwen3.5-9B-MXFP4-GPTQ, scheme swapped to NVFP4) |
| Hardware | 1×B200 (jbom-03), ~2.2 h wall |
| Role in the study | Strong-PTQ baseline row of the stage-3 Qwen3.5-27B 2×2 (W4A4-training vs W4A16-training × MXFP4 vs NVFP4). Companions: Qwen3.5-27B-NVFP4-RTN (weak PTQ) and the NVFP4 QAD arms as they complete. |
# gptq_27b.py NVFP4 (llmcompressor==0.12.0, transformers==5.9.0)
from datasets import load_dataset
from transformers import AutoTokenizer, Qwen3_5ForConditionalGeneration
from llmcompressor import oneshot
from llmcompressor.modifiers.gptq import GPTQModifier
model = Qwen3_5ForConditionalGeneration.from_pretrained("Qwen/Qwen3.5-27B", dtype="bfloat16")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3.5-27B")
ds = load_dataset("json", data_files="calib_512.jsonl", split="train")
recipe = GPTQModifier(
targets="Linear",
scheme="NVFP4",
ignore=["lm_head", "re:.*visual.*", "re:.*mtp.*", "re:.*embed_tokens.*"],
)
oneshot(model=model, tokenizer=tokenizer, dataset=ds, recipe=recipe,
max_seq_length=8192, num_calibration_samples=512)
model.save_pretrained("Qwen3.5-27B-NVFP4-GPTQ", save_compressed=True)
tokenizer.save_pretrained("Qwen3.5-27B-NVFP4-GPTQ")
Serving
vllm serve weili-0234/Qwen3.5-27B-NVFP4-GPTQ # W4A4 (native NVFP4 activations, SM100+)
Serving-KL vs BF16 (2026-07-24)
Mean top-20 token KL vs the BF16 teacher (Qwen/Qwen3.5-27B) over a held-out 256-conversation corpus (818,944 scored tokens), vLLM on 1×B200, W4A4 serving.
| checkpoint | KL vs BF16 | top-1 agree |
|---|---|---|
| NVFP4-RTN (QATFactory export, running-max input scales) | 0.0375 | — |
| this repo (NVFP4-GPTQ, llm-compressor calibrated input scales) | 0.0982 | 0.917 |
Caveat — this artifact does NOT beat RTN on serving-KL for NVFP4. Unlike MXFP4 (where GPTQ recovers ~23% of the RTN gap), the llm-compressor NVFP4 pipeline lands 2.6× worse than the QATFactory RTN export at W4A4 serving. The leading hypothesis is the input-activation global-scale calibration (static observer scales here vs running-max scales in the RTN export) dominating NVFP4's small weight-quantization error; the GPTQ weight update itself may be neutral-to-positive. Treat this checkpoint as a faithful record of the standard llm-compressor GPTQ-NVFP4 recipe at 27B rather than as a strengthened PTQ baseline.
A 7-benchmark evaluation (RULER, AIME'25 avg@4, GPQA-Diamond, MMLU-Pro, HumanEval, IFEval, Pile-10k/WikiText perplexity) against the RTN and QAD checkpoints is in progress and will be added to this card.
- Downloads last month
- 21
Model tree for weili-0234/Qwen3.5-27B-NVFP4-GPTQ
Base model
Qwen/Qwen3.5-27B