Compactbot commited on
Commit
093384d
·
verified ·
1 Parent(s): a2358aa

Add config, tokenizer config, and README

Browse files
Files changed (3) hide show
  1. README.md +79 -0
  2. config.json +16 -0
  3. tokenizer_config.json +8 -0
README.md ADDED
@@ -0,0 +1,79 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ library_name: transformers
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - gpt2
7
+ - tiny
8
+ - tiny-lm
9
+ - small-language-model
10
+ - subword
11
+ - bpe
12
+ - from-scratch
13
+ - tinystories
14
+ - 7m-params
15
+ ---
16
+
17
+ # Subword GPT 7M
18
+
19
+ A 6.95M-parameter GPT-2 style language model trained **from scratch** on [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories) using a custom BPE-8192 tokenizer.
20
+
21
+ ## Architecture
22
+
23
+ | Parameter | Value |
24
+ |-----------|-------|
25
+ | Layers | 6 |
26
+ | Hidden dim | 256 |
27
+ | Heads | 8 |
28
+ | FFN dim | 1024 |
29
+ | Vocab | 8192 (BPE) |
30
+ | Max position | 512 |
31
+ | Tied embeddings | Yes |
32
+ | Biases | No |
33
+ | Norm | RMSNorm |
34
+ | Activation | GELU |
35
+ | **Total params** | **6,950,144** |
36
+
37
+ ## Training
38
+
39
+ - **Data**: TinyStories (~10M BPE tokens after tokenization)
40
+ - **Batch size**: 32 sequences × 512 tokens
41
+ - **Steps**: 3,175 (best checkpoint)
42
+ - **LR schedule**: Cosine decay with warmup
43
+ - **Hardware**: 32-core CPU, ~2 hours
44
+ - **Best val loss**: 3.8398
45
+
46
+ ## Evaluation
47
+
48
+ | Metric | Value |
49
+ |--------|-------|
50
+ | Perplexity (held-out TinyStories, 100×512) | 268.57 |
51
+ | Perplexity (mid-dataset, 50×512) | 302.45 |
52
+
53
+ The held-out perplexity is computed on the last 2M tokens (not seen during training). The gap between train val loss (3.84) and held-out perplexity (268.6) reflects the difficulty of the TinyStories distribution at this model size.
54
+
55
+ ## Why subword?
56
+
57
+ This model is a direct comparison to my earlier [char-gpt-1.2m](https://huggingface.co/Compactbot/char-gpt-1.2m) (character-level, 1.2M params). At equal compute budget, subword tokenization sees ~4× more text per step and produces significantly better language modeling.
58
+
59
+ ## Usage
60
+
61
+ ```python
62
+ from transformers import AutoTokenizer, AutoModelForCausalLM
63
+
64
+ model = AutoModelForCausalLM.from_pretrained("Compactbot/subword-gpt-7m")
65
+ tok = AutoTokenizer.from_pretrained("Compactbot/subword-gpt-7m")
66
+
67
+ text = "Once upon a time, there was a little cat."
68
+ inputs = tok(text, return_tensors="pt")
69
+ outputs = model.generate(**inputs, max_new_tokens=100, do_sample=True, temperature=0.8)
70
+ print(tok.decode(outputs[0], skip_special_tokens=True))
71
+ ```
72
+
73
+ ## Limitations
74
+
75
+ - Trained only on TinyStories (simple English stories for children)
76
+ - 512-token context window
77
+ - Will produce repetitive or incoherent text on out-of-distribution inputs
78
+ - Not a chat model, not instruction-tuned
79
+ - Quality is limited by the 7M parameter budget
config.json ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": ["GPT"],
3
+ "vocab_size": 8192,
4
+ "n_layer": 6,
5
+ "n_head": 8,
6
+ "n_embd": 256,
7
+ "n_inner": 1024,
8
+ "n_positions": 512,
9
+ "tie_word_embeddings": true,
10
+ "bias": false,
11
+ "norm": "rmsnorm",
12
+ "activation": "gelu",
13
+ "torch_dtype": "bfloat16",
14
+ "model_type": "gpt2",
15
+ "transformers_version": "4.40.0"
16
+ }
tokenizer_config.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "tokenizer_class": "GPT2TokenizerFast",
3
+ "model_max_length": 512,
4
+ "pad_token": null,
5
+ "eos_token": "":[[ctrl:endoftext]]",
6
+ "bos_token": "":[[ctrl:endoftext]]",
7
+ "unk_token": null
8
+ }