--- license: apache-2.0 base_model: HuggingFaceTB/SmolLM3-3B tags: - glq - quantization - e8-lattice - rht - ldlq - block-diagonal pipeline_tag: text-generation language: - en --- # SmolLM3-3B-GLQ-block-diagonal-3.5bpw GLQ (E8-lattice) **3.5 bpw** quantization of [HuggingFaceTB/SmolLM3-3B](https://huggingface.co/HuggingFaceTB/SmolLM3-3B), using a **block-diagonal** randomized Hadamard transform — no power-of-2 FHT padding, so the stored size matches the nominal rate (a genuine 3.5 bpw, ~1.8 GB). > **vs the earlier [xv0y5ncu/SmolLM3-3B-GLQ-3.5bpw](https://huggingface.co/xv0y5ncu/SmolLM3-3B-GLQ-3.5bpw):** that checkpoint > zero-padded non-power-of-2 dimensions up to the next power of two for the FHT, which > inflated the *effective* rate to ~5.0 bpw (~2.3 GB). This block-diagonal version stores > only the real weights — smaller, with an accurate bpw label. At equal storage, > block-diagonal also yields higher SQNR, so it is a strict per-bit improvement. ## Quality — lm-evaluation-harness 5-task GLQ and bf16 measured on the **same stack** (vLLM backend, transformers 5.x). | Task | GLQ 3.5 bpw | bf16 | GLQ / bf16 | |---|---|---|---| | arc_easy | 0.7689 | 0.7710 | 99.7% | | hellaswag (acc_norm) | 0.7268 | 0.7573 | 96.0% | | piqa | 0.7709 | 0.7829 | 98.5% | | winogrande | 0.6732 | 0.6638 | 101.4% | | lambada_openai | 0.6532 | 0.6676 | 97.8% | | **average** | **0.7186** | **0.7285** | **98.6%** | Tasks: arc_easy, hellaswag, piqa, winogrande, lambada_openai (`acc_norm` where available, else `acc`). Per-task ratios slightly above 100% are within lm-eval noise. ## Details - Bits per weight: 3.5 bpw average, mixed-precision (2-4 bpw) — block-diagonal, so effective is approximately nominal. - Checkpoint size: ~1.8 GB (bf16 original ~6 GB). - Average SQNR: 16.9 dB. - Calibration: WikiText-2, 128 samples, sequence length 2048. - Method: E8 lattice codebook + randomized Hadamard transform (block-diagonal) + LDLQ. ## Usage ```bash pip install glq ``` ```python import glq.hf_integration # registers GLQ with transformers import torch from transformers import AutoModelForCausalLM, AutoTokenizer repo = "xv0y5ncu/SmolLM3-3B-GLQ-block-diagonal-3.5bpw" tok = AutoTokenizer.from_pretrained(repo) model = AutoModelForCausalLM.from_pretrained(repo, device_map="cuda", dtype=torch.float16) ids = tok("The capital of France is", return_tensors="pt").to("cuda") print(tok.decode(model.generate(**ids, max_new_tokens=20)[0], skip_special_tokens=True)) ``` Serve with vLLM: `vllm serve xv0y5ncu/SmolLM3-3B-GLQ-block-diagonal-3.5bpw --quantization glq`. ## License Apache 2.0. Original model: [HuggingFaceTB/SmolLM3-3B](https://huggingface.co/HuggingFaceTB/SmolLM3-3B). Derivative quantization work; the base model and the GLQ tooling ([github.com/cnygaard/glq](https://github.com/cnygaard/glq)) are both Apache 2.0. --- 🔗 GLQ on GitHub: https://github.com/cnygaard/glq — if you like it, a ⭐ is appreciated.