You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Qwen3.5-9B-MXFP4-GPTQ (baseline)

GPTQ MXFP4 quantization of Qwen/Qwen3.5-9B @ c202236 via llm-compressor 0.12.0 (GPTQModifier, scheme=MXFP4, 512 fixed calibration conversations from openperfectblend-100k think, seed 42, max_seq_length 8192; ignore: lm_head, visual, mtp, embed_tokens). PTQ baseline row of a QATFactory MXFP4 QAD experiment — same serving contract as weili-0234/Qwen3.5-9B-MXFP4-RTN (see that card for vLLM inference instructions).

Results (same fixed harness as the RTN card)

Serving mode held-out KL vs BF16 GSM8K (500) GPQA-Diamond MMLU-Pro (1000)
BF16 base 0 83.80 67.68 77.00
W4A16 (Marlin) 0.0333 84.60 60.61 75.00
W4A4 (FlashInfer, native) 0.1239 77.40 59.09 66.70

Produced with AI assistance (Claude).

Downloads last month
-
Safetensors
Model size
6B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for weili-0234/Qwen3.5-9B-MXFP4-GPTQ

Finetuned
Qwen/Qwen3.5-9B
Quantized
(544)
this model