Commit ·
7366f4d
1
Parent(s): 0396c3b
Fix two factual errors in the v11 card
Browse files1. The corpus is NOT "narrow, template-heavy pinniped text" — it is a mixed
corpus (~88% general FineWeb English / ~12% pinniped, interleaved), as the
training script loads sealglazer_mixed.bin. Measured by sampling 2000 random
128-token windows.
2. The "fix is a mixed corpus" line was self-contradictory (v11 was already the
mixed corpus). The real problem is undertraining (~6.2 tok/param) plus a val
metric that scores a single fixed 256-token window and memorizes it — the
0.038 val is an artifact, not a quality signal. (#4)
- Fix two factual errors in the v11 card (0b4829fc69a02e0febc073edeeaf13401ea40bd0)
README.md
CHANGED
|
@@ -24,17 +24,22 @@ A **from-scratch** subword language model trained to write about pinnipeds
|
|
| 24 |
> seeds it collapses into single-token loops ("like like like…", "is is is…").
|
| 25 |
> It is **not recommended** for real use. It is published here so the requester
|
| 26 |
> can inspect the actual weights and samples, and as an honest record of what
|
| 27 |
-
> happens when a
|
| 28 |
-
>
|
| 29 |
> failure is instability, not a total lack of signal.
|
| 30 |
|
| 31 |
## Why it is degenerate
|
| 32 |
|
| 33 |
-
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
- Lower temperature makes it **worse** (temp 0.3 → 96–98% single-token
|
| 39 |
dominance), the classic sign of a collapsed distribution, not a sampling
|
| 40 |
artifact.
|
|
@@ -60,8 +65,11 @@ This is a model-level collapse, not a bug in the sampler.
|
|
| 60 |
|
| 61 |
## Data
|
| 62 |
|
| 63 |
-
|
| 64 |
-
|
|
|
|
|
|
|
|
|
|
| 65 |
|
| 66 |
## Samples (temp 0.7, top_p 0.9)
|
| 67 |
|
|
@@ -94,8 +102,10 @@ python load_sealglazer.py --prompt "The harbor seal is" --seed 999
|
|
| 94 |
|
| 95 |
## Status
|
| 96 |
|
| 97 |
-
**Public / not recommended for real use.**
|
| 98 |
-
|
| 99 |
-
|
|
|
|
|
|
|
| 100 |
|
| 101 |
_Trained and verified by @Compactbot, 2026-09-22._
|
|
|
|
| 24 |
> seeds it collapses into single-token loops ("like like like…", "is is is…").
|
| 25 |
> It is **not recommended** for real use. It is published here so the requester
|
| 26 |
> can inspect the actual weights and samples, and as an honest record of what
|
| 27 |
+
> happens when a sub-2M model is undertrained on a small corpus. The one
|
| 28 |
+
> coherent seed (999) shows the model *can* form pinniped sentences — the
|
| 29 |
> failure is instability, not a total lack of signal.
|
| 30 |
|
| 31 |
## Why it is degenerate
|
| 32 |
|
| 33 |
+
- **Undertrained for its size.** 1.9M params on ~11.8M tokens is only ~6.2
|
| 34 |
+
tokens/param, well below the several-tens-of-tokens/param a model needs to
|
| 35 |
+
generalize. It memorizes local patterns instead of learning a stable
|
| 36 |
+
distribution, so generation collapses into single-token loops.
|
| 37 |
+
- **The val number is an artifact, not a quality signal.** The training
|
| 38 |
+
harness scores a *single fixed 256-token window* (general English, at the
|
| 39 |
+
train/val boundary). The model memorizes that one window, so "best val 0.038"
|
| 40 |
+
floors around step 2000 and never moves — while general generation is
|
| 41 |
+
degenerate. A near-zero val loss here measures window memorization, not
|
| 42 |
+
capability.
|
| 43 |
- Lower temperature makes it **worse** (temp 0.3 → 96–98% single-token
|
| 44 |
dominance), the classic sign of a collapsed distribution, not a sampling
|
| 45 |
artifact.
|
|
|
|
| 65 |
|
| 66 |
## Data
|
| 67 |
|
| 68 |
+
Mixed corpus: ~11.8M tokens, BPE-8192, from scratch — no base model.
|
| 69 |
+
Composition (measured by sampling 2000 random 128-token windows): **~88%
|
| 70 |
+
general English** (FineWeb-Edu) / **~12% pinniped text**, interleaved. The
|
| 71 |
+
corpus is *not* narrow or template-heavy; the degeneracy is a scale/
|
| 72 |
+
undertraining problem, not a data-narrowness problem.
|
| 73 |
|
| 74 |
## Samples (temp 0.7, top_p 0.9)
|
| 75 |
|
|
|
|
| 102 |
|
| 103 |
## Status
|
| 104 |
|
| 105 |
+
**Public / not recommended for real use.** Note: this model was *already*
|
| 106 |
+
trained on the mixed general+pinniped corpus, so a wider corpus alone will not
|
| 107 |
+
fix it. The levers that matter are **more tokens/param** (train longer or use a
|
| 108 |
+
larger corpus) and a **proper held-out val set** (a fixed 256-token window
|
| 109 |
+
memorizes and misleads). A coherent version needs scale, not just data mixing.
|
| 110 |
|
| 111 |
_Trained and verified by @Compactbot, 2026-09-22._
|