--- license: mit library_name: symbolic-ai tags: - language-model - baseline - symbolic-ai datasets: - HuggingFaceFW/fineweb-edu --- # symbolic-ai-001-lm-baseline The **conventional-pretraining control** for [`jacob-valdez/symbolic-ai-001-3b46de3`](https://huggingface.co/jacob-valdez/symbolic-ai-001-3b46de3): the *identical* 13.97M-parameter architecture, the *identical* 8,000-step budget and optimizer, trained on the FineWeb-Edu corpus component alone. It exists to answer one question: how much of an agent does ordinary next-token pretraining give you, at fixed architecture and compute? **Answer, on held-out seeds:** better language modelling (perplexity **51.9** vs 89.1) and an agent at or below the random policy on every interactive family — gridworld 0.00 success, tool-calling 0.00 solved, arithmetic 0.00 exact, MiniGrid 0.28 vs random 0.40. ## Provenance | | | |---|---| | source repository | `symbolic-ai-001`, snapshot in [the companion repo](https://huggingface.co/jacob-valdez/symbolic-ai-001-3b46de3/tree/main/code) | | source commit | `3b46de309a53f5afc2c5018a1116096de95fbf31` (`3b46de3`) | | adapter version | 0.1.0 | | corpus | `HuggingFaceFW/fineweb-edu`, `sample-10BT`, revision `87f09149ef4734204d70ed1d046ddc9ca3f2b8f9`, shard `sample/10BT/000_00000.parquet` | | split rule | `sha256("symbolic-ai-001\|" + doc_id) < 0.01 → validation` | | training | 8,000 steps × batch 24 × 512 = 98.3M tokens; AdamW, cosine, bf16; 22.2 min on one NVIDIA GB10 | | held-out | val NLL 3.950, perplexity 51.9 | Same tokenizer (`tokenizer.json`), same config shape (`baseline-lm-only.yaml`), same evaluation protocol. Full analysis: `code/docs/symbolic-ai-001.md` in the companion repo. ## License MIT. Trained from random initialization.