SmolLM2 / SmolLM3 - GLQ quantized
SmolLM2 135M/360M and SmolLM3-3B - the small end: CI, demos, and cheap codebook A/Bs.
Text Generation • 1.0B • Updated • 3.95kNote Trellis (TCQ) 4 bpw - the fastest-decoding SmolLM3.
xv0y5ncu/SmolLM3-3B-GLQ-block-diagonal-3.5bpw
Text Generation • 0.9B • Updated • 199Note Block-diagonal RHT, true-to-label footprint.
xv0y5ncu/SmolLM3-3B-GLQ-3.5bpw
Text Generation • 1B • Updated • 348Note Earlier 3.5 bpw quant.
xv0y5ncu/SmolLM3-3B-GLQ-6bpw
Text Generation • 1B • Updated • 11Note 6 bpw.
xv0y5ncu/SmolLM2-360M-Instruct-GLQ-block-diagonal-4bpw
Text Generation • 0.1B • Updated • 184Note 360M, block-diagonal.
xv0y5ncu/SmolLM2-360M-Instruct-GLQ-4bpw
Text Generation • 0.2B • Updated • 426Note 360M, earlier quant.
xv0y5ncu/SmolLM2-135M-Instruct-GLQ-block-diagonal-4bpw
Text Generation • 55.2M • Updated • 197Note 135M, block-diagonal.
xv0y5ncu/SmolLM2-135M-Instruct-GLQ-4bpw
Text Generation • 95.7M • Updated • 295Note 135M, earlier quant.
xv0y5ncu/SmolLM3-3B-trellis-3inst-5bpw
1B • Updated • 13Note Trellis 5 bpw (stacked RVQ 4+1). wikitext2 PPL 9.1402 vs bf16 9.1220. Needs glq >= 0.8.0.
xv0y5ncu/SmolLM3-3B-trellis-3inst-6bpw
1B • Updated • 16Note Trellis 6 bpw (stacked RVQ 4+2). wikitext2 PPL 9.1310 vs bf16 9.1220 - the closest to bf16 of the ladder. Needs glq >= 0.8.0.
xv0y5ncu/SmolLM3-3B-trellis-3inst-3bpw
Text Generation • 0.8B • Updated • 170Note Trellis (3INST) 3 bpw - SmolLM3 ladder rung
xv0y5ncu/SmolLM2-360M-Instruct-GLQ-trellis-3inst-6bpw
0.2B • Updated • 11Note 0.31 GiB · wikitext-2 PPL +0.16% vs bf16
xv0y5ncu/SmolLM2-360M-Instruct-GLQ-trellis-3inst-5bpw
0.1B • Updated • 7Note 0.27 GiB · PPL +0.78% vs bf16
xv0y5ncu/SmolLM2-360M-Instruct-GLQ-trellis-3inst-4bpw
0.1B • Updated • 5Note 0.24 GiB · PPL +2.7% vs bf16
xv0y5ncu/SmolLM2-360M-Instruct-GLQ-trellis-3inst-3bpw
0.1B • Updated • 5Note 0.20 GiB · PPL +11.3% vs bf16 · fastest decode of the ladder