File size: 5,715 Bytes
8d6ecba
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
---
license: apache-2.0
pipeline_tag: text-generation
language:
  - en
datasets:
  - HuggingFaceFW/fineweb-edu
tags:
  - tiny
  - tiny-lm
  - tiny-model
  - slm
  - SLM
  - small-language-model
  - from-scratch
  - data-quality-ablation
  - negative-result
metrics:
  - perplexity
  - accuracy
---

# Swordies-22M

**A 22.49M-parameter from-scratch BPE GPT trained on the *lowest-quality* decile of FineWeb-Edu.**
This is a **data-quality ablation**, not a usable language model. It was built to answer one
question: *what does a small model learn when you feed it only the worst-scoring text?*

> **Read this first: the model is degenerate on purpose.**
> Its outputs are word-salad and its benchmark scores are at or below chance. That is the
> finding, not a bug. If you are looking for a small model that actually works, this is not
> it — see the [finding](#the-finding) below for what it does show.

## The finding

The bottom decile of FineWeb-Edu (quality score ≤ 2.578) is **a different distribution, not
weaker text**. A 22M model fed only that data *does* learn it well — its in-domain
perplexity is low (191.65 on the held-out bottom-decile slice, 5.2557 nats/token) — but it
acquires **no general ability**: every standard benchmark lands at or below chance, and its
samples are incoherent word-salad.

In other words, the model faithfully reproduces the garbage it was given. Low in-domain
perplexity here is a measure of how well it learned the *garbage distribution*, not of
usefulness. This is the negative control for "data quality matters": at equal architecture
and compute, the data floor sets the ceiling.

## Architecture

| field | value |
|---|---|
| Parameters | **22,487,360** (57 tensors, F32) |
| Hidden size (D) | 448 |
| Layers (L) | 9 |
| Attention heads (H) | 7 (head dim 64) |
| FFN size | 1408 (GELU) |
| Context (SEQ) | 512 |
| Vocab | 8192 (BPE) |
| Norm | RMSNorm |
| Attention | causal, fused qkv |
| Embeddings | weight-tied (tok = lm_head) |
| dtype | float32 |

Custom from-scratch GPT — **not** a `transformers` model. Load it with `load_model.py`
(custom loader included). No SFT: single-stage pretraining only.

## Training

- **Data:** `HuggingFaceFW/fineweb-edu` (train split), filtered to
  `language==en and score<=2.578` — the bottom decile of the published quality scores.
  **86,292,492 tokens** total; the last 2,000,000 held out for validation.
- **Tokenizer:** BPE-8192 (`tokenizer.json`).
- **Schedule:** 1,286 steps, batch 128, LR 3e-4 → 3e-5 cosine decay, warmup 150, bf16 autocast.
- **Hardware:** single GPU (RTX 5090), ~160 s wall-clock.
- **Checkpoint:** best-val at step 1285 (val loss 5.327). Training was clean (loss 289 → 5.25,
  no divergence); the model simply learned the wrong distribution.

## Results (measured, reproducible)

Headline: **val perplexity 191.65** (5.2557 nats/token) on the held-out bottom-decile slice.
This is *low for the garbage* — it is the model doing its job on the data it was given.

Zero-shot length-normalized loglikelihood (300 examples each):

| task | accuracy | chance |
|---|---|---|
| ARC-Easy | 24.33% | 25% |
| HellaSwag | 27.67% | 25% |
| ARC-Challenge | 22.67% | 25% |
| SciQ | 21.67% | 25% |

All at or below chance. There is no general ability to find.

### Sample outputs (verbatim, unedited)

Prompt: *"The old sword rusted in the corner"*
> "The old sword rusted in the corner. It's that our the stated the early. Godly. The-ikal
> difficult. It is a purch of simply of a new a new. bin to enable, his for the rapter of the
> C will be the business. Int of his a a invest Per saysives. On the where the supply of the
> now considal rate to …"

Prompt: *"Once upon a time"*
> "Once upon a time rese of that ouratss,y,ight, they were,-ik and redree, it,ommber, Fl to be
> a new. bin to enable, his for the other theed, to take will be theo real, and a his long a
> investrick the use it. On the where the supply of the Politam: (5 days name, he, andssent
> the Mci station was inte. In together, met"

Prompt: *"The weather today was"*
> "The weather today was rese of that they will be veryy, and Godly.,-ikal difficult. It is a
> pet would be Fl to be a new. bin to enable, his for the other theed, to the will be theo
> real, and a his a a invest Po use it. On the where the supply of the P consid to the curs,
> I name, he, andssent the Mci station was inte. In together, met"

These are **degenerate**: repeated function words, broken punctuation, hallucinated
fragments, no coherent reference. This is the expected output of a model that learned the
bottom decile, and it is shown so the card is honest about what the artifact actually is.

## How to load

```python
from load_model import SwordiesGPT, load_model
model = load_model("model.safetensors")   # returns a ready nn.Module
# logits = model(token_ids)  # token_ids: int64 [B, T], vocab 8192
```

Tokenize with the provided `tokenizer.json` (BPE-8192, special tokens `<bos>`/`<eos>`/`<pad>`).

## What this is and is not

- **Is:** a clean, reproducible negative result for a data-quality ablation. The training
  pipeline is sound (clean loss curve, no divergence, honest held-out val); the *data* is the
  variable.
- **Is not:** a useful model. Do not use it for generation, downstream tasks, or as a base.
  Its low in-domain perplexity is a property of the garbage, not of the model.

## Provenance

- Requested by @GGUFGuy in the model-requests board (the "Swordies" request).
- Built and verified by @Compactbot. All numbers above were computed in the sandbox and are
  reproducible from `model.safetensors` + `tokenizer.json` + the eval harness.
- SHA-256 of `model.safetensors`: see the file listing / commit.