--- language: - en license: cc-by-nc-sa-4.0 library_name: transformers pipeline_tag: text-generation tags: - text-generation - causal-lm - custom-architecture - slm - small-language-model - spab datasets: - HuggingFaceFW/fineweb-edu - HuggingFaceTB/cosmopedia - agentlans/high-quality-english-sentences - nampdn-ai/tiny-strange-textbooks - armanc/ScienceQA - nvidia/OpenMathInstruct-2 - microsoft/orca-math-word-problems-200k --- # Tokle-SPAB-3M ## Model Summary Tokle-SPAB-3M is a decoder-only language model trained on 12B tokens, with 2.91M trainable parameters and 8.39M frozen SPAB parameters, for a total of 11.3M parameters. Its main architectural addition is SPAB (Static Pairwise Attention Bias), a frozen table of token-pair association scores built from Pointwise Mutual Information (PMI) over the training corpus and added to the attention logits. For every query-key pair, SPAB hashes the two token IDs into the table, pulls out their PMI value, multiplies it by a learned per-head scale, and adds it to the attention logits before softmax. The bias ignores position and depends only on which tokens are involved, so the model starts training already knowing which tokens tend to co-occur. It only has to learn how much to trust that prior. ## Model Architecture | Parameter | Value | | --- | --- | | Architecture | Custom decoder-only transformer + SPAB | | Layers | 9 | | Hidden size (d_model) | 144 | | Attention heads | 3 | | KV heads (GQA) | 1 (multi-query attention) | | Head dim | 48 | | FFN intermediate size | 432 | | Max sequence length | 512 | | Trainable parameters | 2,908,947 | | Frozen SPAB table | 8,388,608 (float32 buffer) | ## How to use This model uses a custom architecture, so it needs `trust_remote_code=True`. ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "techdotus/Tokle-SPAB-3M" tok = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True).eval() ids = tok("The climate change", return_tensors="pt") with torch.no_grad(): out = model.generate(**ids, max_new_tokens=32, do_sample=False, repetition_penalty=1.3) # greedy print(tok.decode(out[0], skip_special_tokens=True)) ``` ## Benchmark Results All scores are 0-shot acc_norm, using the Open SLM Leaderboard methodology: | Hellaswag | ARC-Easy | ARC-Challenge | PIQA | Arithmark-3 | | --------- | -------- | ------------- | ---- | ----------- | | 27.22% | 34.68% | 24.49% | 54.95% | 41.70% | ## Training Data Details We trained on a curated mixture with a strict cleaning pipeline that also removed topics not useful for a model of this size. | Source | Percentage | | --- | --- | | FineWeb-Edu | 43.1% | | Cosmopedia | 24.3% | | OpenMathInstruct-2 | 13.5% | | Tiny Strange Textbooks | 9.0% | | MegaScience (medicine & biology, custom curated) | 5.0% | | High-Quality English Sentences | 3.0% | | ScienceQA | 1.2% | | Orca-Math Word Problems 200k | 0.9% | | Total | 100% | - **Tokenizer:** all data was tokenized with the model's 5,048-token BPE tokenizer, and 1% was held out for validation. - **Blending:** sources were blended per dataset using the weights above. ## Limitations - **Tiny model:** with ~2.9M trainable parameters, ~8.39M frozen SPAB parameters and 144-dim hidden states, generations are often repetitive, incoherent or factually wrong. The model is a research artifact for studying small-scale LMs, not an assistant. - **Short context:** 512 tokens maximum. RoPE tables are not built beyond that length. - **English only:** trained on English web, educational, synthetic and math text. - **Not instruction-tuned or safety-aligned:** it may reproduce biases present in web data. ## Licenses **Code**: MIT. The modeling code, tokenizer, and training scripts are released under the MIT license. **Weights**: CC BY-NC-SA 4.0. The training data includes MegaScience (CC BY-NC-SA 4.0), so the weights are released under the same terms: attribution required, non-commercial use only, and derivatives (including fine-tunes) must be shared under the same license. ## Citation ```bibtex @misc{tokle2026, title = {{Tokle-SPAB-3M}: Pointwise Mutual Information as an Inductive Bias for Self-Attention}, author = {{Tech.us Team}}, year = {2026}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/techdotus/Tokle-SPAB-3M}} } ```