|
Download h4_fineweb_edu_22x320/README.md from SlayerLab/gollem-v5-ckpts: direct link, hf CLI and curl.
- Browser
- Download file 1.65 kB
-
https://huggingface.co/SlayerLab/gollem-v5-ckpts/resolve/fb0ca1ef0ab6ce541bf4f7f26d4960b528a5b83f/h4_fineweb_edu_22x320/README.md
- Command line
-
hf download hf://SlayerLab/gollem-v5-ckpts@fb0ca1ef0ab6ce541bf4f7f26d4960b528a5b83f/h4_fineweb_edu_22x320/README.md
-
curl -L -o README.md https://huggingface.co/SlayerLab/gollem-v5-ckpts/resolve/fb0ca1ef0ab6ce541bf4f7f26d4960b528a5b83f/h4_fineweb_edu_22x320/README.md
1.65 kB
h4_fineweb_edu_22x320 — H4: deep-narrow 32M (22 layers × 320) trained from scratch on FineWeb-Edu
Why it exists: tests whether a deeper, narrower shape at the same parameter count improves grammatical knowledge (BLiMP) without losing ARC-Easy. It uses the same data pool, recipe and token budget as Z1 arm Z (z1_Z_fineweb_edu/, 15 layers × 384); only the model shape differs.
Setup
- Model: 31.0M parameters, 22 layers, d_model 320, 5 heads (head dim 64), RoPE, SwiGLU, RMSNorm, QK-norm, value residual; BPE tokenizer with 12,288 tokens (
tokenizer.jsonin the repository root). - Recipe: identical to
z1_Z_fineweb_edu/: Muon (hidden 2-D weights) + AdamW, learning rate 6e-4 with 2,000 warmup steps and cosine decay to 6e-5, batch 32 × 1024 tokens, 150,000 steps (4.9B tokens), seed 1337. Training windows are drawn without replacement (each window at most once), less than one epoch. - Data: the same 5.0B-token FineWeb-Edu pool as
z1_Z_fineweb_edu/(ODC-BY; three sources deduplicated by document hash), scanned against WikiText-2 and ARC test/validation with normalized 13-gram and short-question matching; matching documents were removed. - Checkpoints: every 30,000 steps (about 1B tokens); the last one is step 150,000.
- Status: research checkpoints, not a leaderboard submission. The comparison is with Z1 arm Z (seeds 1337 and 1338); the architecture and trainer differ from our published 32M, so their numbers are not directly comparable.
- Format: PyTorch checkpoint dict with
model,opt,step,config;train_gpt_ref.pyin the repository root rebuilds the model fromconfig.