Kahnn nano (~1.8M) โ CPU Chinchilla run + lifelong teach
Neuroscience-inspired non-Transformer model from AFKmoney/kahnn.
Files
| File | What |
|---|---|
ckpt_final.pt |
Nano after ~37M tokens CPU pretrain (2026-09-09) |
ckpt_teach.pt |
Same + taught fact ยซ La capitale du Canada est Ottawa. ยป |
Code
All training / teach / forget code lives in the GitHub repo (not duplicated here):
train_universal.py,teach.py,generate.py- Docs:
docs/UNIVERSAL_TRAINING.md,docs/LIFELONG_LEARNING.md,docs/CPU_RUN_LOG.md
git clone https://github.com/AFKmoney/kahnn
cd kahnn
pip install -r requirements.txt
# download weights from this Hub repo into ./runs/
python generate.py --checkpoint ./runs/nano_base/ckpt_final.pt --prompt "Once upon a time" --device cpu
python teach.py probe --text "capitale du Canada" --resume ./runs/teach/ckpt_teach.pt --config nano
Honest metrics (CPU box, 2026-09-09)
- ~36.99M tokens, ~1362 tok/s average at end, CE loss still ~10.5 (generation still weak)
- Teach exact fact โ memory coherence ~1.0; partial probe ~0.43
- Prefer lifelong engram memory over fluent generation at this size/budget
License
See GitHub repo.
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support