Model Card for Model ID

This is a mixed BF16-INT8 AWQ layer quantization, with working MTP (speculative decoding) via llmcompressor.

Model Details

📖 Model Description:

The "NM" in the name refers to "neuralmagic/LLM_compression_calibration" dataset used for this quant.

Fixed chat_template with "froggeric/Qwen-Fixed-Chat-Templates"

Working MTP with VLLM flag:

--speculative-config '{"method":"mtp","num_speculative_tokens":2}'

Tested with VLLM 0.19.1 and transformers 5.6.2

Recommended flags:

--enable-auto-tool-choice
--reasoning-parser qwen3
--tool-call-parser qwen3_xml

🏆 Ranking: Qwopus3.5-9B-v3 Series (just my benchmarks):

Rank Model / Dataset HumanEval (Code) ↑ Winogrande (Logic) ↑ HellaSwag (Context) ↑ WikiText (PPL) ↓ Verdict
1st AWQ-NM (NeuralMagic) 0.6768 0.7395 0.7820 9.6056 Best All-Rounder
2nd AWQ-UC (Ultrachat) 0.6768 0.7466 0.7814 9.6069 Best for Chat/Reasoning
3rd AWQ-CK (CyanKiwi) 0.6707 0.7427 0.7813 9.6054 Highest Fidelity
4th Base (BF16) 0.6890 0.7427 0.7813 9.6042 Reference (Slow)

🔍 Key Takeaways:

  • NeuralMagic (NM): The only version to surpass the base model in context understanding (HellaSwag).
  • Ultrachat200k (UC): Achieved the highest score in logical reasoning (Winogrande).
  • CyanKiwi (CK): Maintained the lowest perplexity (WikiText), showing minimal knowledge loss during quantization.
  • Efficiency: AWQ versions are ~35% faster than the Base model with negligible accuracy trade-offs.

🙏 Acknowledgements:

Downloads last month
11
Safetensors
Model size
10B params
Tensor type
I64
·
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Arien0/Qwopus3.5-9B-v3-AWQ-BF16-INT8-NM-MTP

Finetuned
Qwen/Qwen3.5-9B
Quantized
(19)
this model

Dataset used to train Arien0/Qwopus3.5-9B-v3-AWQ-BF16-INT8-NM-MTP