symbolic-ai-001-lm-baseline

The conventional-pretraining control for jacob-valdez/symbolic-ai-001-3b46de3: the identical 13.97M-parameter architecture, the identical 8,000-step budget and optimizer, trained on the FineWeb-Edu corpus component alone.

It exists to answer one question: how much of an agent does ordinary next-token pretraining give you, at fixed architecture and compute?

Answer, on held-out seeds: better language modelling (perplexity 51.9 vs 89.1) and an agent at or below the random policy on every interactive family — gridworld 0.00 success, tool-calling 0.00 solved, arithmetic 0.00 exact, MiniGrid 0.28 vs random 0.40.

Provenance

source repository symbolic-ai-001, snapshot in the companion repo
source commit 3b46de309a53f5afc2c5018a1116096de95fbf31 (3b46de3)
adapter version 0.1.0
corpus HuggingFaceFW/fineweb-edu, sample-10BT, revision 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9, shard sample/10BT/000_00000.parquet
split rule sha256("symbolic-ai-001|" + doc_id) < 0.01 → validation
training 8,000 steps × batch 24 × 512 = 98.3M tokens; AdamW, cosine, bf16; 22.2 min on one NVIDIA GB10
held-out val NLL 3.950, perplexity 51.9

Same tokenizer (tokenizer.json), same config shape (baseline-lm-only.yaml), same evaluation protocol. Full analysis: code/docs/symbolic-ai-001.md in the companion repo.

License

MIT. Trained from random initialization.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train jacob-valdez/symbolic-ai-001-lm-baseline-3b46de3