base_model:
- Qwen/Qwen3.5-2B
license: apache-2.0
Vishva007/Qwen3.5-2B-W4A16-AutoRound
This is a W4A16 (4-bit weight, 16-bit activation) quantized version of Qwen/Qwen3.5-2B, produced using AutoRound β Intel's sign gradient descent based quantization method designed for production-grade accuracy retention. MTP Enabled model quantization
Quantization Details
| Parameter | Value |
|---|---|
| Method | AutoRound (W4A16) |
| Group Size | 32 |
| Symmetric | Yes |
| Iterations | 1000 |
| Calibration Samples | 512 |
| Sequence Length | 4096 |
| Torch Compile | Enabled |
Key Notes
- High accuracy configuration β 1000 iterations with 512 calibration samples targets production-grade quality with minimal degradation from the base model.
- W4A16 β Weights are quantized to 4-bit integers; activations remain in FP16 for inference stability.
- ~50% memory reduction compared to the FP16 base model, enabling deployment on consumer and mid-range GPUs.
- MTP (Multi-Token Prediction) Enabled β Supports speculative decoding for faster inference.
MTP / Speculative Decoding
This model supports Multi-Token Prediction (MTP) for improved inference throughput using speculative decoding.
When serving with compatible backends (e.g., vLLM), enable MTP using:
--speculative_config '{"method":"mtp","num_speculative_tokens":3}'
Notes
num_speculative_tokens=1is a stable default for balancing speed and accuracy.- You can experiment with higher values for better throughput, depending on your hardware and latency requirements.
Usage
This model is compatible with transformers and backends that support AutoRound GPTQ-format weights (e.g., vLLM, SGLang, AutoGPTQ). For full model details, architecture, and capabilities, refer to the base model page.
π Deploy on RunPod
One-click launch environments pre-configured with PyTorch, CUDA, and dependencies for fine-tuning or quantization.
π Need GPU compute? Sign up via RunPod and get $5β$500 in free credits when you add your first $10.