Compactbot commited on
Commit
e0d0017
·
verified ·
1 Parent(s): 6529b56

Update card to 4000-step checkpoint: ppl 268.57→55.50, val loss 3.8398→3.7659

Browse files
Files changed (1) hide show
  1. README.md +5 -6
README.md CHANGED
@@ -38,19 +38,18 @@ A 6.95M-parameter GPT-2 style language model trained **from scratch** on [TinySt
38
 
39
  - **Data**: TinyStories (~10M BPE tokens after tokenization)
40
  - **Batch size**: 32 sequences × 512 tokens
41
- - **Steps**: 3,175 (best checkpoint)
42
  - **LR schedule**: Cosine decay with warmup
43
  - **Hardware**: 32-core CPU, ~2 hours
44
- - **Best val loss**: 3.8398
45
 
46
  ## Evaluation
47
 
48
  | Metric | Value |
49
  |--------|-------|
50
- | Perplexity (held-out TinyStories, 100×512) | 268.57 |
51
- | Perplexity (mid-dataset, 50×512) | 302.45 |
52
 
53
- The held-out perplexity is computed on the last 2M tokens (not seen during training). The gap between train val loss (3.84) and held-out perplexity (268.6) reflects the difficulty of the TinyStories distribution at this model size.
54
 
55
  ## Why subword?
56
 
@@ -76,4 +75,4 @@ print(tok.decode(outputs[0], skip_special_tokens=True))
76
  - 512-token context window
77
  - Will produce repetitive or incoherent text on out-of-distribution inputs
78
  - Not a chat model, not instruction-tuned
79
- - Quality is limited by the 7M parameter budget
 
38
 
39
  - **Data**: TinyStories (~10M BPE tokens after tokenization)
40
  - **Batch size**: 32 sequences × 512 tokens
41
+ - **Steps**: 4,000 (best checkpoint)
42
  - **LR schedule**: Cosine decay with warmup
43
  - **Hardware**: 32-core CPU, ~2 hours
44
+ - **Best val loss**: 3.7659
45
 
46
  ## Evaluation
47
 
48
  | Metric | Value |
49
  |--------|-------|
50
+ | Perplexity (held-out TinyStories, 100×512) | 55.50 |
 
51
 
52
+ The held-out perplexity is computed on the last 2M tokens (not seen during training). The gap between train val loss (3.77) and held-out perplexity (55.5) reflects the difficulty of the TinyStories distribution at this model size.
53
 
54
  ## Why subword?
55
 
 
75
  - 512-token context window
76
  - Will produce repetitive or incoherent text on out-of-distribution inputs
77
  - Not a chat model, not instruction-tuned
78
+ - Quality is limited by the 7M parameter budget