Transformers
Safetensors
English
charlm
tiny
tiny-lm
small-language-model
sub-1m
char-level
from-scratch
nanoGPT
TinyStories
Instructions to use Compactbot/char-gpt-1.2m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Compactbot/char-gpt-1.2m with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import CharGPT model = CharGPT.from_pretrained("Compactbot/char-gpt-1.2m", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Add model card: architecture, data, exact param count, honest assessment of quality.
Browse files
README.md
ADDED
|
@@ -0,0 +1,91 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language: en
|
| 3 |
+
license: apache-2.0
|
| 4 |
+
library_name: transformers
|
| 5 |
+
tags:
|
| 6 |
+
- tiny
|
| 7 |
+
- tiny-lm
|
| 8 |
+
- small-language-model
|
| 9 |
+
- sub-1m
|
| 10 |
+
- char-level
|
| 11 |
+
- from-scratch
|
| 12 |
+
- nanoGPT
|
| 13 |
+
- TinyStories
|
| 14 |
+
base_model: []
|
| 15 |
+
---
|
| 16 |
+
|
| 17 |
+
# Char-GPT 1.2M
|
| 18 |
+
|
| 19 |
+
A tiny **character-level** causal transformer trained **from scratch** on
|
| 20 |
+
[TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories). A small,
|
| 21 |
+
honest reference build — the point is a model whose card matches its artifact
|
| 22 |
+
exactly, not a competitive checkpoint.
|
| 23 |
+
|
| 24 |
+
## Parameters (exact)
|
| 25 |
+
|
| 26 |
+
**1,216,000 parameters, untied head.**
|
| 27 |
+
|
| 28 |
+
| module | params |
|
| 29 |
+
|---|---|
|
| 30 |
+
| `transformer.wte` (65×128) | 8,320 |
|
| 31 |
+
| `transformer.wpe` (128×128) | 16,384 |
|
| 32 |
+
| 6 × attention (qkv + proj, bias-free) | 614,400 |
|
| 33 |
+
| 6 × FFN (4×, bias-free) | 552,960 |
|
| 34 |
+
| 6 × 2 LayerNorm (affine) | 1,536 |
|
| 35 |
+
| `ln_f` (128) | 256 |
|
| 36 |
+
| `lm_head` (65×128, **separate / untied**) | 8,320 |
|
| 37 |
+
| **total** | **1,216,000** |
|
| 38 |
+
|
| 39 |
+
> The head is **not** weight-tied: the checkpoint stores two distinct 65×128
|
| 40 |
+
> tensors (`transformer.wte.weight` and `lm_head.weight`), and `model.py` never
|
| 41 |
+
> assigns one to the other. `config.json` therefore says
|
| 42 |
+
> `tie_word_embeddings: false`. (If the head were tied the count would be
|
| 43 |
+
> 1,207,680.)
|
| 44 |
+
|
| 45 |
+
## Architecture
|
| 46 |
+
|
| 47 |
+
nanoGPT-style GPT-2, all bias-free except LayerNorm:
|
| 48 |
+
- `n_layer=6`, `n_head=4`, `n_embd=128`, FFN = 4× = 512
|
| 49 |
+
- `vocab_size=65` (printable ASCII + newline), `block_size=128`
|
| 50 |
+
- RoPE: none (learned positional embedding `wpe`)
|
| 51 |
+
|
| 52 |
+
## Training
|
| 53 |
+
|
| 54 |
+
- **Data:** `roneneldan/TinyStories` (train split), first ~1.0M characters,
|
| 55 |
+
90/5/5 train/val/test split by character.
|
| 56 |
+
- **Steps:** 1,500, batch 32 × seq 128, AdamW (lr 6e-4, cosine, warmup),
|
| 57 |
+
grad-clip 1.0, float32, CPU (16 threads). ~3.5 min.
|
| 58 |
+
- **Seed:** 42.
|
| 59 |
+
|
| 60 |
+
## Quality — what it is and is not
|
| 61 |
+
|
| 62 |
+
Held-out perplexities (from training log):
|
| 63 |
+
|
| 64 |
+
| split | loss | perplexity |
|
| 65 |
+
|---|---|---|
|
| 66 |
+
| val | 1.9046 | **6.72** |
|
| 67 |
+
| test | 1.9473 | **7.01** |
|
| 68 |
+
|
| 69 |
+
It captures TinyStories' surface style (short declarative sentences, simple
|
| 70 |
+
vocabulary, character names) but it is a **1.2M-parameter model on ~1M
|
| 71 |
+
characters** — it does not grasp meaning, it repeats and drifts, and it will
|
| 72 |
+
produce the kind of plausible-looking-but-nonsense text in `sample.txt`.
|
| 73 |
+
Treat it as a working toy / reference architecture, not a useful language model.
|
| 74 |
+
|
| 75 |
+
## Files
|
| 76 |
+
|
| 77 |
+
- `model.safetensors` — 4,869,136 B (53 tensors, F32)
|
| 78 |
+
- `model.py` — `CharGPT` + `from_config`
|
| 79 |
+
- `config.json`, `tokenizer_config.json` (char vocab)
|
| 80 |
+
- `sample.txt` — 240-char greedy-ish sample
|
| 81 |
+
- `LICENSE` — Apache-2.0
|
| 82 |
+
|
| 83 |
+
## Reproduce
|
| 84 |
+
|
| 85 |
+
```python
|
| 86 |
+
import torch, json
|
| 87 |
+
from model import from_config
|
| 88 |
+
cfg = json.load(open("config.json"))
|
| 89 |
+
m = from_config(cfg)
|
| 90 |
+
print(sum(p.numel() for p in m.parameters())) # 1216000
|
| 91 |
+
```
|