Compactbot commited on
Commit
012bb0c
·
verified ·
1 Parent(s): e0d0017

Add measured zero-shot benchmark table (ARC-Easy/Challenge, HellaSwag, SciQ) + note on held-out ppl variance

Browse files
Files changed (1) hide show
  1. README.md +25 -4
README.md CHANGED
@@ -11,6 +11,7 @@ tags:
11
  - bpe
12
  - from-scratch
13
  - tinystories
 
14
  - 7m-params
15
  ---
16
 
@@ -45,15 +46,30 @@ A 6.95M-parameter GPT-2 style language model trained **from scratch** on [TinySt
45
 
46
  ## Evaluation
47
 
 
 
48
  | Metric | Value |
49
  |--------|-------|
50
- | Perplexity (held-out TinyStories, 100×512) | 55.50 |
 
 
 
 
 
 
51
 
52
- The held-out perplexity is computed on the last 2M tokens (not seen during training). The gap between train val loss (3.77) and held-out perplexity (55.5) reflects the difficulty of the TinyStories distribution at this model size.
 
 
 
 
 
 
 
53
 
54
  ## Why subword?
55
 
56
- This model is a direct comparison to my earlier [char-gpt-1.2m](https://huggingface.co/Compactbot/char-gpt-1.2m) (character-level, 1.2M params). At equal compute budget, subword tokenization sees ~4× more text per step and produces significantly better language modeling.
57
 
58
  ## Usage
59
 
@@ -75,4 +91,9 @@ print(tok.decode(outputs[0], skip_special_tokens=True))
75
  - 512-token context window
76
  - Will produce repetitive or incoherent text on out-of-distribution inputs
77
  - Not a chat model, not instruction-tuned
78
- - Quality is limited by the 7M parameter budget
 
 
 
 
 
 
11
  - bpe
12
  - from-scratch
13
  - tinystories
14
+ - charlm
15
  - 7m-params
16
  ---
17
 
 
46
 
47
  ## Evaluation
48
 
49
+ ### Held-out perplexity
50
+
51
  | Metric | Value |
52
  |--------|-------|
53
+ | Perplexity (held-out TinyStories, 100×512) | 48.43 |
54
+
55
+ The held-out perplexity is computed on the last 2M tokens (not seen during training). Note that this figure has meaningful **sample variance**: a 100×512 draw gives 48.43 (seed 123) while the first 20×512 slice gives 55.39 — both are honest draws from the same distribution, and the spread (not a bug) reflects how uneven the TinyStories difficulty is at this model size. The gap between train val loss (3.77) and held-out perplexity reflects the difficulty of the TinyStories distribution at this model size.
56
+
57
+ ### Zero-shot benchmarks
58
+
59
+ Measured with length-normalized loglikelihood scoring (each answer choice scored as a continuation of the prompt; argmax of mean per-token logprob vs. gold). 400 examples per task, 32-core CPU.
60
 
61
+ | Task | Split | Accuracy | Chance (4-choice) |
62
+ |------|-------|----------|-------------------|
63
+ | ARC-Easy | test | 23.5% | 25% |
64
+ | ARC-Challenge | test | 19.5% | 25% |
65
+ | HellaSwag | validation | 23.75% | 25% |
66
+ | SciQ | test | 23.0% | 25% |
67
+
68
+ All four tasks sit **at or below the 25% four-choice chance level**. This is the honest expectation for a 7M-parameter model trained only on TinyStories: the corpus carries no general reasoning or commonsense signal, so the model cannot do better than chance on these out-of-distribution tasks. These numbers are reported so the card states what the model is *not* good at, not just what it is.
69
 
70
  ## Why subword?
71
 
72
+ This model is a direct comparison to my earlier [char-gpt-1.2m](https://huggingface.co/Compactbot/char-gpt-1.2m) (character-level, 1.2M params). At equal compute budget, subword tokenization sees ~4× more text per step and produces significantly better language modeling. This is the "subword beats character" result.
73
 
74
  ## Usage
75
 
 
91
  - 512-token context window
92
  - Will produce repetitive or incoherent text on out-of-distribution inputs
93
  - Not a chat model, not instruction-tuned
94
+ - Near-chance on general reasoning/commonsense benchmarks (see table above) — no general knowledge signal in the training data
95
+ - Quality is limited by the 7M parameter budget
96
+
97
+ ## Reproduction
98
+
99
+ Training script: see `generation.py` for inference. The training code is available in the CompactAI workspace. Benchmark harness: `eval_bench.py` (loglikelihood scoring) and `eval_validate2.py` (held-out perplexity).