Qwen3.5-0.8B-EXL3-3.0bpw

Qwen3.5-0.8B quantized to 3.0 bits-per-weight with the EXL3 (bitshift trellis) format, using the 3inst codebook.

Quantized with exllamav3.

Model Description

This is an EXL3-format quantization of Qwen/Qwen3.5-0.8B, the lightweight vision-language variant of the Qwen3.5 family. See the base model card for architecture details, intended use, and limitations.

Property Value
Base model Qwen/Qwen3.5-0.8B
Quantization format EXL3 (QTIP-style bitshift trellis)
Bits per weight 3.0 (head: 6.0)
Codebook 3inst
Parameters ~0.8B
License Apache 2.0

Quantization Details

  • Format: EXL3 — bitshift trellis coding with procedural codebook decode.
  • Calibration: 256 rows x 2048 columns from a general-domain text corpus mix (C4 / Wikipedia / code / technical).
  • Applied output-channel scales: always.

Evaluation

Metric fp16 3.0bpw EXL3 Delta
Perplexity (Wiki) 11.21 12.24 +9.2%
KL div (fp16 || 3bit) 0.101
Top-1 accuracy 0.525 0.512 −2.5pp

Usage

This repository is intended for use with EXL3-capable inference runtimes (ExLlamaV3). It is not a standalone Transformers-format model.

# ExLlamaV3
python chat.py -m BlivionIaG/Qwen3.5-0.8B-exl3-3.0bpw

Credits

This is a community quantization; the base model's capabilities and limitations apply unchanged.

Downloads last month
18
Safetensors
Model size
0.4B params
Tensor type
F32
·
BF16
·
F16
·
I16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BlivionIaG/Qwen3.5-0.8B-exl3-3.0bpw

Quantized
(234)
this model