--- license: mit library_name: transformers pipeline_tag: text-generation tags: - gpt2 - tiny - tiny-lm - small-language-model - subword - bpe - from-scratch - tinystories - 7m-params --- # Subword GPT 7M A 6.95M-parameter GPT-2 style language model trained **from scratch** on [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories) using a custom BPE-8192 tokenizer. ## Architecture | Parameter | Value | |-----------|-------| | Layers | 6 | | Hidden dim | 256 | | Heads | 8 | | FFN dim | 1024 | | Vocab | 8192 (BPE) | | Max position | 512 | | Tied embeddings | Yes | | Biases | No | | Norm | RMSNorm | | Activation | GELU | | **Total params** | **6,950,144** | ## Training - **Data**: TinyStories (~10M BPE tokens after tokenization) - **Batch size**: 32 sequences × 512 tokens - **Steps**: 4,000 (best checkpoint) - **LR schedule**: Cosine decay with warmup - **Hardware**: 32-core CPU, ~2 hours - **Best val loss**: 3.7659 ## Evaluation | Metric | Value | |--------|-------| | Perplexity (held-out TinyStories, 100×512) | 55.50 | The held-out perplexity is computed on the last 2M tokens (not seen during training). The gap between train val loss (3.77) and held-out perplexity (55.5) reflects the difficulty of the TinyStories distribution at this model size. ## Why subword? This model is a direct comparison to my earlier [char-gpt-1.2m](https://huggingface.co/Compactbot/char-gpt-1.2m) (character-level, 1.2M params). At equal compute budget, subword tokenization sees ~4× more text per step and produces significantly better language modeling. ## Usage ```python from transformers import AutoTokenizer, AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Compactbot/subword-gpt-7m") tok = AutoTokenizer.from_pretrained("Compactbot/subword-gpt-7m") text = "Once upon a time, there was a little cat." inputs = tok(text, return_tensors="pt") outputs = model.generate(**inputs, max_new_tokens=100, do_sample=True, temperature=0.8) print(tok.decode(outputs[0], skip_special_tokens=True)) ``` ## Limitations - Trained only on TinyStories (simple English stories for children) - 512-token context window - Will produce repetitive or incoherent text on out-of-distribution inputs - Not a chat model, not instruction-tuned - Quality is limited by the 7M parameter budget