You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Qwen3.5-27B-NVFP4-RTN-weight-only

Data-free RTN (round-to-nearest) NVFP4 weight-only (W4A16) quantization of Qwen/Qwen3.5-27B in vLLM compressed-tensors format: input_activations is null in the quantization config and no input_global_scale tensors exist — activations run bf16. Weight bytes are identical to the W4A4-schema sibling Qwen3.5-27B-NVFP4-RTN; the two differ only in the activation contract.

How this checkpoint was produced

Producing repo tonyzhang-together/QATFactory branch weili/qwen35-27b @ 096f2ae
Base model Qwen/Qwen3.5-27B (dense qwen3_5, bf16)
Training None. Step-0 / PTQ baseline: the quantized weights are the base bf16 weights rounded to the lattice by the same exporter used for all QAD checkpoints in this series (round-to-nearest, RTN). At 9B this construction was verified byte-identical to llm-compressor model_free_ptq.
Role in the study Day-0 activation-gap gate + PTQ baseline row of the stage-3 Qwen3.5-27B 2x2 (W4A4-training vs W4A16-training, MXFP4 vs NVFP4); companion QAD checkpoints appear under weili-0234/Qwen3.5-27B-*-QAD-* as the arms complete.
Evaluation Serving-KL vs the BF16 teacher + benchmark rows are being measured (2026-07-22); the results table in this card is updated as rows land.

Exporter invocation

python scripts/export_nvfp4_vllm.py --source <calib ckpt (global_step=20)> \
  --model-assets Qwen3.5-27B --output <this repo> --device cuda:0 --weight-only

Serving

vllm serve weili-0234/Qwen3.5-27B-NVFP4-RTN-weight-only   # W4A16

Evaluation results (2026-07-23)

vLLM on 1×B200, identical harness across rows. GSM8K = 500-question subset; GPQA-Diamond is a single seed (treat ±2–3 pts as noise); MMLU-Pro full. KL = mean top-20 token KL divergence vs the BF16 teacher (Qwen/Qwen3.5-27B) over a held-out 256-conversation corpus (818,944 scored tokens). Throughput = generation tok/s during the GSM8K run. These are the RTN (no training) baselines of the W4A4-vs-W4A16 training study; QAD-trained checkpoints follow separately.

serving mode KL vs BF16 GSM8K GPQA-D MMLU-Pro gen tok/s
W4A16 (weight-only NVFP4) 0.0234 91.8 81.8 86.0 3243
BF16 teacher 0 93.6 83.3 87.1 —

For native W4A4 serving use the a4-schema export (weili-0234/Qwen3.5-27B-NVFP4-RTN): KL 0.0375, GSM8K 92.6, GPQA-D 77.8, MMLU-Pro 85.6.

Downloads last month
-
Safetensors
Model size
17B params
Tensor type
BF16
·
F32
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for weili-0234/Qwen3.5-27B-NVFP4-RTN-weight-only

Base model

Qwen/Qwen3.5-27B
Quantized
(230)
this model