Maggio33's picture
READMEs for Z4 (FineWeb-Edu score >= 4) and H4 (deep-narrow 22x320)
6207b8f verified
|
Raw History Blame
1.65 kB

z4_fineweb_edu_ge4 — Z4: 32M trained from scratch on FineWeb-Edu with educational score >= 4

Why it exists: tests whether a stricter educational filter builds more ARC-Easy than the standard FineWeb-Edu threshold. It uses the same model, recipe and token budget as Z1 arm Z (z1_Z_fineweb_edu/, standard threshold); only the training data differs.

Setup

  • Model: 31.3M parameters, 15 layers, d_model 384, 6 heads (head dim 64), RoPE, SwiGLU, RMSNorm, QK-norm, value residual; BPE tokenizer with 12,288 tokens (tokenizer.json in the repository root).
  • Recipe: identical to z1_Z_fineweb_edu/: Muon (hidden 2-D weights) + AdamW, learning rate 6e-4 with 2,000 warmup steps and cosine decay to 6e-5, batch 32 × 1024 tokens, 150,000 steps (4.9B tokens), seed 1337. Training windows are drawn without replacement (each window at most once), less than one epoch.
  • Data: 5.0B-token pool of FineWeb-Edu (ODC-BY) documents with educational score >= 4 from the first 50 files of the 100BT sample (pinned revision), deduplicated by document hash, scanned against WikiText-2 and ARC test/validation with normalized 13-gram and short-question matching; matching documents were removed.
  • Checkpoints: every 30,000 steps (about 1B tokens); the last one is step 150,000.
  • Status: research checkpoints, not a leaderboard submission. The comparison is with Z1 arm Z; the architecture and trainer differ from our published 32M, so their numbers are not directly comparable.
  • Format: PyTorch checkpoint dict with model, opt, step, config; train_gpt_ref.py in the repository root rebuilds the model from config.