Maggio33 commited on
Commit
8793375
Β·
verified Β·
1 Parent(s): 544ffdf

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +102 -18
README.md CHANGED
@@ -1,25 +1,109 @@
1
- # GoLLeM-v5 β€” modele (backup lokalny)
 
 
 
 
 
 
 
 
 
 
 
 
2
 
3
- Backup ckptΓ³w treningowych GoLLeM-v5 (scan tokenΓ³w 16M). ΕΉrΓ³dΕ‚o: RunPod RTX 5090.
4
- Redundancja: pod (oryginaΕ‚) + HF `SlayerLab/minimal-en-corpus-5b:ckpts/` (durable) + ten katalog (lokalny).
5
 
6
- ## Modele (scan tokenΓ³w, single-factor, BPE-12k, L6/d408/h6, block1024, batch64, AdamW lr6e-4β†’6e-5, seed1337)
 
 
 
7
 
8
- | plik | tokeny | epoki | BLiMP | ARC-Easy | WikiText-2 BPB | board-pozycja |
9
- |------|--------|-------|-------|----------|----------------|---------------|
10
- | run_bpe16m_c_ckpt.pt | 3.2B | 1.19 | 71.54 | 39.39 | 1.2161 | #16/74 |
11
- | run_bpe16m_6b_d_ckpt.pt | 6B | 2.2 | 73.29 | 41.04 | 1.1943 | #11/74 |
12
- | run_bpe16m_10b_e_ckpt.pt| 10B | 3.7 | 73.43 | 41.67 | 1.1815 | (liczy Latarnik) |
13
 
14
- **Wniosek scanu:** BLiMP saturuje ~6B na 16M (Ξ”6β†’10=+0.14); ARC dalej roΕ›nie @10B (knee>10B); BPB monotoniczny.
15
 
16
- - `tokenizer.json` β€” BPE-12k (vocab 12288), ZALOCKOWANY (werdykt 3/3 vs byte-baseline: BLiMP+2Οƒ ∧ BPB-Ξ”0.238 ∧ board-#16). WspΓ³lny dla wszystkich modeli v5.
17
- - Config eval: `byte_lm_eval_bpe.py --tokenizer tokenizer.json --ckpt <model> --tasks wikitext,blimp,arc_easy --n-head 6`.
18
- - Korpus treningowy: `SlayerLab/minimal-en-corpus-5b` (2.70B unique BPE-tok, FineWeb-Edu, ARC-targeted, decontam'd).
19
 
20
- ## Kolejne (dojdΔ… po treningu)
21
- - 32M-baseline (L6/d576/h9, izolacja-rozmiaru) + lewary single-factor (RoPE/low-rank/Muon/depth).
22
- - 150M flagowiec (+KARD-distill z Qwena).
 
 
 
 
23
 
24
- Prowieniencja peΕ‚na: labvault `21_09_GoLLeM-v5-Skalowanie-Glint/` (Stan.md + artifacts + S2-Wnioski).
25
- Data backupu: 2026-09-22.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ tags:
6
+ - tiny-lm
7
+ - babylm
8
+ - language-model
9
+ - from-scratch
10
+ - gpt
11
+ library_name: pytorch
12
+ pipeline_tag: text-generation
13
+ ---
14
 
15
+ # GoLLeM-v5 β€” English Tiny-LM Checkpoints
 
16
 
17
+ Training checkpoints for **GoLLeM-v5**, a family of small English language models
18
+ (β‰ˆ17M–31M parameters) trained **from scratch** on a curated, decontaminated English
19
+ corpus. Built for the **sub-100M efficiency regime** and evaluated on the tiny-ML
20
+ leaderboard tasks: **BLiMP**, **ARC-Easy**, and **WikiText-2**.
21
 
22
+ Trained on a single RunPod **RTX 5090**. Redundant backups: this HF repo (durable) +
23
+ the training pod (origin) + local disk.
 
 
 
24
 
25
+ ## Architecture
26
 
27
+ All 16M checkpoints share one decoder-only GPT (`GPT-ref`):
 
 
28
 
29
+ | | |
30
+ |---|---|
31
+ | Layers / width / heads | **6 / 408 / 6** (β‰ˆ17.4M params) |
32
+ | Context length | 1024 |
33
+ | Vocabulary | 12288 (BPE-12k) |
34
+ | Optimizer | AdamW, lr 6e-4 β†’ 6e-5 (cosine) |
35
+ | Precision / seed | bf16 / 1337 |
36
 
37
+ The 32M baseline is the same recipe at **L6 / d576 / h9 (β‰ˆ31.4M params)**.
38
+
39
+ ## Checkpoints
40
+
41
+ ### Token-scaling scan (16M, single-factor: identical arch/hypers, only the token budget changes)
42
+
43
+ | file | tokens | epochs | BLiMP | ARC-Easy | WikiText-2 (BPB) |
44
+ |------|-------:|-------:|------:|---------:|-----------------:|
45
+ | `bpe16m_3.2B/ckpt.pt` | 3.2B | 1.2 | 71.54 | 39.39 | 1.2161 |
46
+ | `bpe16m_6B/ckpt.pt` | 6B | 2.2 | 73.29 | 41.04 | 1.1943 |
47
+ | `bpe16m_10B/ckpt.pt` | 10B | 3.7 | 73.43 | 41.67 | 1.1815 |
48
+
49
+ Metrics above are from `lm-evaluation-harness` (BLiMP full, ARC-Easy, WikiText-2 bits-per-byte).
50
+
51
+ **Finding:** BLiMP saturates around ~6B tokens for this 16M model (Ξ” 6β†’10B = +0.14);
52
+ ARC-Easy keeps improving at 10B; BPB decreases monotonically.
53
+
54
+ ### 32M size-isolation baseline
55
+
56
+ | file | params | tokens |
57
+ |------|-------:|-------:|
58
+ | `bpe32m_baseline/ckpt.pt` | 31.4M (L6/d576/h9) | 10B |
59
+
60
+ ### Leaderboard protocol (Glint) results
61
+
62
+ The public leaderboard uses a stricter evaluation protocol (256-token clip, no BOS,
63
+ train-split BLiMP, bare-prompt ARC-Easy raw accuracy, byte-normalized WikiText). Under
64
+ that protocol our numbers are lower than the `lm-eval` numbers above (offset β‰ˆ 3.1pp
65
+ BLiMP / 2.2pp ARC) β€” **always cite the Glint numbers for the board**:
66
+
67
+ | model | tokens | BLiMP | ARC-Easy | efficiency rank |
68
+ |-------|-------:|------:|---------:|:---------------:|
69
+ | 16M | 10B | 70.36 | 39.52 | #18 / 74 |
70
+ | 32M | 10B | 70.08 | 42.59 | #20 / 74 |
71
+
72
+ The 32M model ranks *below* the 16M model on efficiency: the leaderboard's size bonus
73
+ favors smaller models, and the extra ARC gain does not offset the reduced bonus β€” hence
74
+ the crown effort stays at the 16M size.
75
+
76
+ ## Tokenizer
77
+
78
+ `tokenizer.json` β€” BPE, vocab 12288 ("BPE-12k"), shared by every v5 model. Locked after a
79
+ 3/3 verdict against a byte-level baseline (BLiMP +2Οƒ ∧ WikiText-2 BPB Ξ”0.238 ∧ improved
80
+ leaderboard efficiency).
81
+
82
+ ## Training data
83
+
84
+ - Base corpus: **`SlayerLab/minimal-en-corpus-5b`** β€” 5.40B unique BPE-12k tokens, a
85
+ 15-source English mixture (FineWeb-Edu, DCLM, code, math, books, QA, science),
86
+ decontaminated against the benchmark test sets.
87
+ - An **expanded 8.29B-token** corpus (base + additional FineWeb-Edu-100BT + OpenStax
88
+ science, all deduplicated and decontaminated) backs the in-progress 16M crown run.
89
+
90
+ ## Usage
91
+
92
+ ```python
93
+ import torch
94
+ ckpt = torch.load("bpe16m_10B/ckpt.pt", map_location="cpu")
95
+ sd = ckpt["model"] if "model" in ckpt else ckpt
96
+ # GPT-ref: L6 d408 h6, vocab 12288, block 1024.
97
+ # Model class + forward: train_gpt_ref.py. Trim logits to vocab 12288 (no padded vocab).
98
+ ```
99
+
100
+ ## Status / roadmap
101
+
102
+ - βœ… 16M token-scaling scan (3.2 / 6 / 10B) and 32M size baseline.
103
+ - πŸ”„ 16M @ expanded 8.29B corpus (16B-token "crown" run) β€” in progress.
104
+ - ⏭ ARC-focused knowledge distillation (KARD) to lift ARC-Easy.
105
+
106
+ ## Provenance
107
+
108
+ Full experiment log and artifacts in the project vault
109
+ (`21_09_GoLLeM-v5-Skalowanie-Glint/`). Backup date: 2026-09-22.