Qwen3.6-35B-A3B-EXL3-5.0bpw

EXL3 quantized version of Qwen/Qwen3.6-35B-A3B at 5.0 bits per weight. Mixture-of-Experts architecture.

Quantization Details

Parameter Value
Source model Qwen/Qwen3.6-35B-A3B
Quantization method EXL3
EXL3 version v0.0.30
Bits per weight 5.0
Codebook mcg
Calibration 250 rows x 2048 columns
Architecture MoE (256 experts, 8 active)

Hardware Used

Component Specification
CPU Intel Core i7-9800X @ 3.80GHz (8C/16T)
Motherboard ASUS WS X299 SAGE
RAM 32GB DDR4-2666
GPU(s) 6x NVIDIA GeForce RTX 3070 8GB
Storage MSI M480 PRO 1TB NVMe
OS Ubuntu 24.04.4 LTS
Driver NVIDIA 580.173.02

Orchestration: Quantization jobs were orchestrated from a dedicated control node (Intel i9-7920X, 62GB RAM, Quadro P5000 + RTX 3070 Ti + 2x RTX 3080) over a dedicated 10Gb interconnect to the compute node above.

File Sizes

File Size
model-00001-of-00003.safetensors 8.45 GB
model-00002-of-00003.safetensors 8.50 GB
model-00003-of-00003.safetensors 6.71 GB
Total 23.70 GB

Notes

  • Custom quantization of the Qwen3.6 MoE variant - not available from other sources at this precision
  • Uses mcg codebook (optimal for MoE architectures)
  • Clean quantization with no hardware errors
Downloads last month
5
Safetensors
Model size
12B params
Tensor type
BF16
·
F16
·
I16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support