|
Download final_64m_14x576/README.md from SlayerLab/gollem-v5-ckpts: direct link, hf CLI and curl.
- Browser
- Download file 2.06 kB
-
https://huggingface.co/SlayerLab/gollem-v5-ckpts/resolve/0da09be998d4e11cb747939d94b2ec6a62f5a541/final_64m_14x576/README.md
- Command line
-
hf download hf://SlayerLab/gollem-v5-ckpts@0da09be998d4e11cb747939d94b2ec6a62f5a541/final_64m_14x576/README.md
-
curl -L -o README.md https://huggingface.co/SlayerLab/gollem-v5-ckpts/resolve/0da09be998d4e11cb747939d94b2ec6a62f5a541/final_64m_14x576/README.md
2.06 kB
final_64m_14x576 — GoLLeM-v5 64M final (14 layers × 576)
Why it exists: a final model of the GoLLeM-v5 series. It is trained from scratch on the ARC-MIX corpus with the recipe of our 64M/128M models and a trainer that draws training windows without accidental repetition; the data composition was chosen by a pre-registered comparison against FineWeb-Edu at 32M (a tie, so the existing corpus was kept).
Setup
- Model: 62.9M parameters, 14 layers, d_model 576, 9 heads (head dim 64), RoPE (theta 100,000), SwiGLU (FFN multiplier 2.667), RMSNorm, QK-norm, value residual; BPE tokenizer with 12,288 tokens (
tokenizer.jsonin the repository root), context 1024. - Recipe: Muon (hidden 2-D weights) + AdamW, peak learning rate 6e-4 (Muon 0.02), 2,000 warmup steps, cosine decay to 6e-5, batch 32 × 1024 tokens, 760,000 steps (24.9B tokens, about 2.65 passes over the pool), seed 1337. Each pass over the pool uses a new permutation of training windows.
- Data: ARC-MIX, 9.39B tokens after scanning against WikiText-2 and ARC test/validation with normalized 13-gram and short-question matching (matching documents were removed); documents containing the marker ‘CC BY-NC-SA’ were also removed. See the root card of this repository for the composition of ARC-MIX. Attributions: FineWeb-Edu (HuggingFaceFW/fineweb-edu, ODC-BY 1.0); OpenStax textbooks (CC-BY 4.0, © Rice University, openstax.org; titles and editions listed in OPENSTAX_ATTRIBUTION.md in this folder); minimal-en-corpus-5b (SlayerLab/minimal-en-corpus-5b; see its card for component licences).
- Checkpoints: every 20,000 steps. The result is the final checkpoint (step 760,000); intermediate checkpoints are public but are not used for selection or for reporting.
- Status: training in progress. The result will be measured with the leaderboard's evaluation protocol on the final checkpoint.
- Format: PyTorch checkpoint dict with
model,opt,step,config;train_gpt_ref.pyin the repository root rebuilds the model fromconfig.