Compactbot commited on
Commit
97c4125
·
verified ·
1 Parent(s): ea61a00

Add model card: architecture, data, exact param count, honest assessment of quality.

Browse files
Files changed (1) hide show
  1. README.md +91 -0
README.md ADDED
@@ -0,0 +1,91 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language: en
3
+ license: apache-2.0
4
+ library_name: transformers
5
+ tags:
6
+ - tiny
7
+ - tiny-lm
8
+ - small-language-model
9
+ - sub-1m
10
+ - char-level
11
+ - from-scratch
12
+ - nanoGPT
13
+ - TinyStories
14
+ base_model: []
15
+ ---
16
+
17
+ # Char-GPT 1.2M
18
+
19
+ A tiny **character-level** causal transformer trained **from scratch** on
20
+ [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories). A small,
21
+ honest reference build — the point is a model whose card matches its artifact
22
+ exactly, not a competitive checkpoint.
23
+
24
+ ## Parameters (exact)
25
+
26
+ **1,216,000 parameters, untied head.**
27
+
28
+ | module | params |
29
+ |---|---|
30
+ | `transformer.wte` (65×128) | 8,320 |
31
+ | `transformer.wpe` (128×128) | 16,384 |
32
+ | 6 × attention (qkv + proj, bias-free) | 614,400 |
33
+ | 6 × FFN (4×, bias-free) | 552,960 |
34
+ | 6 × 2 LayerNorm (affine) | 1,536 |
35
+ | `ln_f` (128) | 256 |
36
+ | `lm_head` (65×128, **separate / untied**) | 8,320 |
37
+ | **total** | **1,216,000** |
38
+
39
+ > The head is **not** weight-tied: the checkpoint stores two distinct 65×128
40
+ > tensors (`transformer.wte.weight` and `lm_head.weight`), and `model.py` never
41
+ > assigns one to the other. `config.json` therefore says
42
+ > `tie_word_embeddings: false`. (If the head were tied the count would be
43
+ > 1,207,680.)
44
+
45
+ ## Architecture
46
+
47
+ nanoGPT-style GPT-2, all bias-free except LayerNorm:
48
+ - `n_layer=6`, `n_head=4`, `n_embd=128`, FFN = 4× = 512
49
+ - `vocab_size=65` (printable ASCII + newline), `block_size=128`
50
+ - RoPE: none (learned positional embedding `wpe`)
51
+
52
+ ## Training
53
+
54
+ - **Data:** `roneneldan/TinyStories` (train split), first ~1.0M characters,
55
+ 90/5/5 train/val/test split by character.
56
+ - **Steps:** 1,500, batch 32 × seq 128, AdamW (lr 6e-4, cosine, warmup),
57
+ grad-clip 1.0, float32, CPU (16 threads). ~3.5 min.
58
+ - **Seed:** 42.
59
+
60
+ ## Quality — what it is and is not
61
+
62
+ Held-out perplexities (from training log):
63
+
64
+ | split | loss | perplexity |
65
+ |---|---|---|
66
+ | val | 1.9046 | **6.72** |
67
+ | test | 1.9473 | **7.01** |
68
+
69
+ It captures TinyStories' surface style (short declarative sentences, simple
70
+ vocabulary, character names) but it is a **1.2M-parameter model on ~1M
71
+ characters** — it does not grasp meaning, it repeats and drifts, and it will
72
+ produce the kind of plausible-looking-but-nonsense text in `sample.txt`.
73
+ Treat it as a working toy / reference architecture, not a useful language model.
74
+
75
+ ## Files
76
+
77
+ - `model.safetensors` — 4,869,136 B (53 tensors, F32)
78
+ - `model.py` — `CharGPT` + `from_config`
79
+ - `config.json`, `tokenizer_config.json` (char vocab)
80
+ - `sample.txt` — 240-char greedy-ish sample
81
+ - `LICENSE` — Apache-2.0
82
+
83
+ ## Reproduce
84
+
85
+ ```python
86
+ import torch, json
87
+ from model import from_config
88
+ cfg = json.load(open("config.json"))
89
+ m = from_config(cfg)
90
+ print(sum(p.numel() for p in m.parameters())) # 1216000
91
+ ```