Qwen3.5-9B-MXFP4-GPTQ (baseline)
GPTQ MXFP4 quantization of Qwen/Qwen3.5-9B @ c202236 via llm-compressor 0.12.0 (GPTQModifier, scheme=MXFP4, 512 fixed calibration conversations from openperfectblend-100k think, seed 42, max_seq_length 8192; ignore: lm_head, visual, mtp, embed_tokens). PTQ baseline row of a QATFactory MXFP4 QAD experiment — same serving contract as weili-0234/Qwen3.5-9B-MXFP4-RTN (see that card for vLLM inference instructions).
Results (same fixed harness as the RTN card)
| Serving mode | held-out KL vs BF16 | GSM8K (500) | GPQA-Diamond | MMLU-Pro (1000) |
|---|---|---|---|---|
| BF16 base | 0 | 83.80 | 67.68 | 77.00 |
| W4A16 (Marlin) | 0.0333 | 84.60 | 60.61 | 75.00 |
| W4A4 (FlashInfer, native) | 0.1239 | 77.40 | 59.09 | 66.70 |
Produced with AI assistance (Claude).
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support