READMEs for Z4 (FineWeb-Edu score >= 4) and H4 (deep-narrow 22x320)
Browse files- h4_fineweb_edu_22x320/README.md +12 -0
- z4_fineweb_edu_ge4/README.md +12 -0
h4_fineweb_edu_22x320/README.md
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# h4_fineweb_edu_22x320 — H4: deep-narrow 32M (22 layers × 320) trained from scratch on FineWeb-Edu
|
| 2 |
+
|
| 3 |
+
**Why it exists:** tests whether a deeper, narrower shape at the same parameter count improves grammatical knowledge (BLiMP) without losing ARC-Easy. It uses the same data pool, recipe and token budget as Z1 arm Z (`z1_Z_fineweb_edu/`, 15 layers × 384); only the model shape differs.
|
| 4 |
+
|
| 5 |
+
## Setup
|
| 6 |
+
|
| 7 |
+
- **Model:** 31.0M parameters, 22 layers, d_model 320, 5 heads (head dim 64), RoPE, SwiGLU, RMSNorm, QK-norm, value residual; BPE tokenizer with 12,288 tokens (`tokenizer.json` in the repository root).
|
| 8 |
+
- **Recipe:** identical to `z1_Z_fineweb_edu/`: Muon (hidden 2-D weights) + AdamW, learning rate 6e-4 with 2,000 warmup steps and cosine decay to 6e-5, batch 32 × 1024 tokens, 150,000 steps (4.9B tokens), seed 1337. Training windows are drawn without replacement (each window at most once), less than one epoch.
|
| 9 |
+
- **Data:** the same 5.0B-token FineWeb-Edu pool as `z1_Z_fineweb_edu/` (ODC-BY; three sources deduplicated by document hash), scanned against WikiText-2 and ARC test/validation with normalized 13-gram and short-question matching; matching documents were removed.
|
| 10 |
+
- **Checkpoints:** every 30,000 steps (about 1B tokens); the last one is step 150,000.
|
| 11 |
+
- **Status:** research checkpoints, **not a leaderboard submission**. The comparison is with Z1 arm Z (seeds 1337 and 1338); the architecture and trainer differ from our published 32M, so their numbers are not directly comparable.
|
| 12 |
+
- **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.
|
z4_fineweb_edu_ge4/README.md
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# z4_fineweb_edu_ge4 — Z4: 32M trained from scratch on FineWeb-Edu with educational score >= 4
|
| 2 |
+
|
| 3 |
+
**Why it exists:** tests whether a stricter educational filter builds more ARC-Easy than the standard FineWeb-Edu threshold. It uses the same model, recipe and token budget as Z1 arm Z (`z1_Z_fineweb_edu/`, standard threshold); only the training data differs.
|
| 4 |
+
|
| 5 |
+
## Setup
|
| 6 |
+
|
| 7 |
+
- **Model:** 31.3M parameters, 15 layers, d_model 384, 6 heads (head dim 64), RoPE, SwiGLU, RMSNorm, QK-norm, value residual; BPE tokenizer with 12,288 tokens (`tokenizer.json` in the repository root).
|
| 8 |
+
- **Recipe:** identical to `z1_Z_fineweb_edu/`: Muon (hidden 2-D weights) + AdamW, learning rate 6e-4 with 2,000 warmup steps and cosine decay to 6e-5, batch 32 × 1024 tokens, 150,000 steps (4.9B tokens), seed 1337. Training windows are drawn without replacement (each window at most once), less than one epoch.
|
| 9 |
+
- **Data:** 5.0B-token pool of FineWeb-Edu (ODC-BY) documents with educational score >= 4 from the first 50 files of the 100BT sample (pinned revision), deduplicated by document hash, scanned against WikiText-2 and ARC test/validation with normalized 13-gram and short-question matching; matching documents were removed.
|
| 10 |
+
- **Checkpoints:** every 30,000 steps (about 1B tokens); the last one is step 150,000.
|
| 11 |
+
- **Status:** research checkpoints, **not a leaderboard submission**. The comparison is with Z1 arm Z; the architecture and trainer differ from our published 32M, so their numbers are not directly comparable.
|
| 12 |
+
- **Format:** PyTorch checkpoint dict with `model`, `opt`, `step`, `config`; `train_gpt_ref.py` in the repository root rebuilds the model from `config`.
|