Compactbot commited on
Commit
7366f4d
·
1 Parent(s): 0396c3b

Fix two factual errors in the v11 card

Browse files

1. The corpus is NOT "narrow, template-heavy pinniped text" — it is a mixed
corpus (~88% general FineWeb English / ~12% pinniped, interleaved), as the
training script loads sealglazer_mixed.bin. Measured by sampling 2000 random
128-token windows.
2. The "fix is a mixed corpus" line was self-contradictory (v11 was already the
mixed corpus). The real problem is undertraining (~6.2 tok/param) plus a val
metric that scores a single fixed 256-token window and memorizes it — the
0.038 val is an artifact, not a quality signal. (#4)


- Fix two factual errors in the v11 card (0b4829fc69a02e0febc073edeeaf13401ea40bd0)

Files changed (1) hide show
  1. README.md +22 -12
README.md CHANGED
@@ -24,17 +24,22 @@ A **from-scratch** subword language model trained to write about pinnipeds
24
  > seeds it collapses into single-token loops ("like like like…", "is is is…").
25
  > It is **not recommended** for real use. It is published here so the requester
26
  > can inspect the actual weights and samples, and as an honest record of what
27
- > happens when a tiny model is trained only on a narrow template corpus. The
28
- > one coherent seed (999) shows the model *can* form pinniped sentences — the
29
  > failure is instability, not a total lack of signal.
30
 
31
  ## Why it is degenerate
32
 
33
- - Trained on a **narrow, template-heavy corpus** (~11.8M tokens of pinniped
34
- text, 60+ recurring sentence templates). A 1.9M model memorizes those
35
- templates rather than learning to vary them.
36
- - Validation loss is near-zero (best **0.038** nats), which on a tiny template
37
- corpus is the signature of **memorization**, not generalization.
 
 
 
 
 
38
  - Lower temperature makes it **worse** (temp 0.3 → 96–98% single-token
39
  dominance), the classic sign of a collapsed distribution, not a sampling
40
  artifact.
@@ -60,8 +65,11 @@ This is a model-level collapse, not a bug in the sampler.
60
 
61
  ## Data
62
 
63
- Pinniped corpus: ~11.8M tokens (general pinniped text + ~60 recurring
64
- sentence templates), BPE-8192. From scratch — no base model.
 
 
 
65
 
66
  ## Samples (temp 0.7, top_p 0.9)
67
 
@@ -94,8 +102,10 @@ python load_sealglazer.py --prompt "The harbor seal is" --seed 999
94
 
95
  ## Status
96
 
97
- **Public / not recommended for real use.** If a coherent version is wanted, the
98
- fix is a mixed corpus (more general text + pinniped) at lower temperature and
99
- more steps — the collapse here is a data/size problem, not an architecture bug.
 
 
100
 
101
  _Trained and verified by @Compactbot, 2026-09-22._
 
24
  > seeds it collapses into single-token loops ("like like like…", "is is is…").
25
  > It is **not recommended** for real use. It is published here so the requester
26
  > can inspect the actual weights and samples, and as an honest record of what
27
+ > happens when a sub-2M model is undertrained on a small corpus. The one
28
+ > coherent seed (999) shows the model *can* form pinniped sentences — the
29
  > failure is instability, not a total lack of signal.
30
 
31
  ## Why it is degenerate
32
 
33
+ - **Undertrained for its size.** 1.9M params on ~11.8M tokens is only ~6.2
34
+ tokens/param, well below the several-tens-of-tokens/param a model needs to
35
+ generalize. It memorizes local patterns instead of learning a stable
36
+ distribution, so generation collapses into single-token loops.
37
+ - **The val number is an artifact, not a quality signal.** The training
38
+ harness scores a *single fixed 256-token window* (general English, at the
39
+ train/val boundary). The model memorizes that one window, so "best val 0.038"
40
+ floors around step 2000 and never moves — while general generation is
41
+ degenerate. A near-zero val loss here measures window memorization, not
42
+ capability.
43
  - Lower temperature makes it **worse** (temp 0.3 → 96–98% single-token
44
  dominance), the classic sign of a collapsed distribution, not a sampling
45
  artifact.
 
65
 
66
  ## Data
67
 
68
+ Mixed corpus: ~11.8M tokens, BPE-8192, from scratch — no base model.
69
+ Composition (measured by sampling 2000 random 128-token windows): **~88%
70
+ general English** (FineWeb-Edu) / **~12% pinniped text**, interleaved. The
71
+ corpus is *not* narrow or template-heavy; the degeneracy is a scale/
72
+ undertraining problem, not a data-narrowness problem.
73
 
74
  ## Samples (temp 0.7, top_p 0.9)
75
 
 
102
 
103
  ## Status
104
 
105
+ **Public / not recommended for real use.** Note: this model was *already*
106
+ trained on the mixed general+pinniped corpus, so a wider corpus alone will not
107
+ fix it. The levers that matter are **more tokens/param** (train longer or use a
108
+ larger corpus) and a **proper held-out val set** (a fixed 256-token window
109
+ memorizes and misleads). A coherent version needs scale, not just data mixing.
110
 
111
  _Trained and verified by @Compactbot, 2026-09-22._