symbolic-ai-001-lm-baseline
The conventional-pretraining control for
jacob-valdez/symbolic-ai-001-3b46de3:
the identical 13.97M-parameter architecture, the identical 8,000-step budget
and optimizer, trained on the FineWeb-Edu corpus component alone.
It exists to answer one question: how much of an agent does ordinary next-token pretraining give you, at fixed architecture and compute?
Answer, on held-out seeds: better language modelling (perplexity 51.9 vs 89.1) and an agent at or below the random policy on every interactive family — gridworld 0.00 success, tool-calling 0.00 solved, arithmetic 0.00 exact, MiniGrid 0.28 vs random 0.40.
Provenance
| source repository | symbolic-ai-001, snapshot in the companion repo |
| source commit | 3b46de309a53f5afc2c5018a1116096de95fbf31 (3b46de3) |
| adapter version | 0.1.0 |
| corpus | HuggingFaceFW/fineweb-edu, sample-10BT, revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9, shard sample/10BT/000_00000.parquet |
| split rule | sha256("symbolic-ai-001|" + doc_id) < 0.01 → validation |
| training | 8,000 steps × batch 24 × 512 = 98.3M tokens; AdamW, cosine, bf16; 22.2 min on one NVIDIA GB10 |
| held-out | val NLL 3.950, perplexity 51.9 |
Same tokenizer (tokenizer.json), same config shape (baseline-lm-only.yaml),
same evaluation protocol. Full analysis: code/docs/symbolic-ai-001.md in the
companion repo.
License
MIT. Trained from random initialization.