SmolLM3-3B-GLQ-block-diagonal-3.5bpw

GLQ (E8-lattice) 3.5 bpw quantization of HuggingFaceTB/SmolLM3-3B, using a block-diagonal randomized Hadamard transform — no power-of-2 FHT padding, so the stored size matches the nominal rate (a genuine 3.5 bpw, ~1.8 GB).

vs the earlier xv0y5ncu/SmolLM3-3B-GLQ-3.5bpw: that checkpoint zero-padded non-power-of-2 dimensions up to the next power of two for the FHT, which inflated the effective rate to 5.0 bpw (2.3 GB). This block-diagonal version stores only the real weights — smaller, with an accurate bpw label. At equal storage, block-diagonal also yields higher SQNR, so it is a strict per-bit improvement.

Quality — lm-evaluation-harness 5-task

GLQ and bf16 measured on the same stack (vLLM backend, transformers 5.x).

Task GLQ 3.5 bpw bf16 GLQ / bf16
arc_easy 0.7689 0.7710 99.7%
hellaswag (acc_norm) 0.7268 0.7573 96.0%
piqa 0.7709 0.7829 98.5%
winogrande 0.6732 0.6638 101.4%
lambada_openai 0.6532 0.6676 97.8%
average 0.7186 0.7285 98.6%

Tasks: arc_easy, hellaswag, piqa, winogrande, lambada_openai (acc_norm where available, else acc). Per-task ratios slightly above 100% are within lm-eval noise.

Details

  • Bits per weight: 3.5 bpw average, mixed-precision (2-4 bpw) — block-diagonal, so effective is approximately nominal.
  • Checkpoint size: ~1.8 GB (bf16 original ~6 GB).
  • Average SQNR: 16.9 dB.
  • Calibration: WikiText-2, 128 samples, sequence length 2048.
  • Method: E8 lattice codebook + randomized Hadamard transform (block-diagonal) + LDLQ.

Usage

pip install glq
import glq.hf_integration  # registers GLQ with transformers
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "xv0y5ncu/SmolLM3-3B-GLQ-block-diagonal-3.5bpw"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, device_map="cuda", dtype=torch.float16)
ids = tok("The capital of France is", return_tensors="pt").to("cuda")
print(tok.decode(model.generate(**ids, max_new_tokens=20)[0], skip_special_tokens=True))

Serve with vLLM: vllm serve xv0y5ncu/SmolLM3-3B-GLQ-block-diagonal-3.5bpw --quantization glq.

License

Apache 2.0. Original model: HuggingFaceTB/SmolLM3-3B. Derivative quantization work; the base model and the GLQ tooling (github.com/cnygaard/glq) are both Apache 2.0.


🔗 GLQ on GitHub: https://github.com/cnygaard/glq — if you like it, a ⭐ is appreciated.

Downloads last month
202
Safetensors
Model size
0.9B params
Tensor type
BF16
·
F16
·
I16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for xv0y5ncu/SmolLM3-3B-GLQ-block-diagonal-3.5bpw

Quantized
(110)
this model

Collections including xv0y5ncu/SmolLM3-3B-GLQ-block-diagonal-3.5bpw