--- license: apache-2.0 library_name: pytorch language: - en pipeline_tag: text-generation tags: - transformer - byte-level - causal-lm - attention - rope - decoder-only - tiny-stories - pretrained-backbone - micro-language-model - sub-1m-parameters model_name: MicroT-test1-50K-TinyStories datasets: - roneneldan/TinyStories metrics: - perplexity ---
MicroMixer-4 Logo # MicroT-test1-50K-TinyStories Parameters Architecture Pretrain

Micro Transformer β€” Reference Baseline
Vanilla Attention β€’ RoPE β€’ Byte-Level β€’ Decoder-Only
[![GitHub](https://img.shields.io/badge/GitHub-MicroMixer--4-blue?style=for-the-badge&logo=github&color=%23007BFF)](https://github.com/llaa33219/MicroMixer-4)
## πŸ“‹ Overview **MicroT-test1-50K-TinyStories** is a **49,888-parameter** vanilla decoder-only **transformer** β€” multi-head causal self-attention with RoPE β€” **pretrained on [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories)** as part of the MicroMixer-4 **dataset-efficiency comparison study** ([analysis](https://github.com/llaa33219/MicroMixer-4/blob/main/DATASET_COMPARISON_ANALYSIS.md)). TinyStories is a corpus of ~200K short synthetic children's stories (GPT-3.5/4-generated, constrained grammar and vocabulary). ⚠️ This repo is **pretrain-only** β€” it is **NOT FMSP-fine-tuned**. TinyStories contains free-flowing prose with no `User:`/`Assistant:` dialogue markers, so the FMSP answer-only cross-entropy recipe does not apply (there is no answer region to isolate). What you get is the **raw pretrained backbone** (seed 42, V76 recipe, 3 epochs). It continues stories; it does not answer questions or follow instructions. It is the registered attention-based reference baseline β€” a deliberately boring standard 2018–2020 transformer recipe (no flash attention, no SwiGLU, no ALiBi, no QKNorm, no sliding window, no MQA/GQA, no MoE). The 50K preset reproduces the champion recipe at its budget.
## πŸ—οΈ Architecture
```mermaid graph TD A[Byte Input] --> B[Embed 256β†’32] B --> C[Transformer Block Γ— 3] C --> D[RMSNorm] D --> E[LM Head Tied with Embed] E --> F[Byte Output] subgraph "Transformer Block (pre-norm)" X[Input 32] --> N1[RMSNorm] N1 --> AT["MHA 2 heads Γ— d_head 16
RoPE ΞΈ=10000 on q,k Β· causal SDPA"] AT --> R1[+ residual] R1 --> N2[RMSNorm] N2 --> MLP["GELU MLP 32β†’152β†’32"] MLP --> R2[+ residual] end style A fill:#007BFF,color:#fff style F fill:#00D620,color:#fff style AT fill:#FF6600,color:#fff ```
### Model Configuration
Parameter Value
Hidden Dimension (d_model)32
Attention Heads2 (d_head = 16 at every size)
Number of Blocks3
FFN Hidden152
Position EncodingRoPE ΞΈ=10000 on q/k only (non-persistent buffers)
AttentionCausal MHA via F.scaled_dot_product_attention(is_causal=True)
ActivationGELU
BiasesNone β€” no bias parameters anywhere
NormalizationRMSNorm (pre-norm)
Max Sequence Length1024
Vocabulary Size256 (byte-level)
Output HeadTied with input embedding
### Core Components ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ Transformer Block (Γ—3) β”‚ β”‚ h = h + MHA(RMSNorm(h)) # RoPE q/k, causalβ”‚ β”‚ h = h + MLP(RMSNorm(h)) # GELU dβ†’ffnβ†’d β”‚ β”‚ no biases, no flash, no tricks β€” vanilla β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ ``` The d_head=16 contract is hard-asserted across all six sizes so that attention-head behavior is comparable at every budget and never confounds the memorization measurements. --- ## 🎯 Generation Examples
**Story-continuation probes** (raw text β€” no `User:/Assistant:` prefix, since TinyStories is a plain LM; greedy decoding, `repetition_penalty=1.2`, `no_repeat_ngram_size=4`, `max_new_tokens=200`): ``` [Prompt] Once upon a time, there was a little girl named Lily. [Output] She loved to play with her friends with her friends. One day, she went to the store to play wit… ``` drifts β€” starts as story but falls into a repetitive loop ``` [Prompt] One day, a boy named Tom found a big red balloon. [Output] They liked to play with his friends and saw a big box. They were so happy and saw a big box and… ``` coherent β€” on-topic story continuation in the TinyStories register ``` [Prompt] The sun was shining and the birds were singing. [Output] They were so happy to the store to the store to play with his friends. They were so happy to th… ``` drifts β€” starts as story but falls into a repetitive loop ``` [Prompt] A little dog ran into the garden and saw [Output] a big box. The bird was so happy to the store to play with his friends. They were so happy and … ``` drifts β€” starts as story but falls into a repetitive loop
--- ## πŸ“Š Results
### Pretraining (TinyStories, V76 recipe, 3 epochs) | Metric | 1 ep | 2 ep | 3 ep | |--------|------|------|------| | Val PPL | 2.85 | 2.76 | **2.60** | AdamW lr 3e-3 Β· WSD (warmup 500) Β· wd 0.01 Β· bs 16 Β· seq 1024 Β· seed 42 Β· plain CE on non-pad bytes. ### MicroT-test1 TinyStories family (pretrain-only, all sizes) | Size | Params | 3ep Val PPL | |------|--------|------------| | 1M | 996,736 | **1.69** | | 500K | 498,528 | **1.79** | | 300K | 297,680 | **1.90** | | 100K | 97,872 | **2.28** | | **50K** | 49,888 | **2.60** | | 10K | 9,808 | **4.31** | Pretrain-only family β€” FMSP-based axes (chatter fluency, full-988 EM, q-relevance, OOD, unanswerable fabrication) are N/A: TinyStories has no `User:/Assistant:` markers, so the answer-only-CE recipe does not apply.
--- ## πŸ“š Training Data
1. **Pretraining**: [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories) β€” synthetic short stories generated by GPT-3.5/4 with a constrained vocabulary and simple grammar, ~200K stories sampled, flattened to 1024-byte sequences, 3 epochs. No `User:/Assistant:` dialogue structure.
--- ## πŸ”§ Usage ### Files in this repository - `epoch_{0,1,2}.safetensors` β€” per-epoch pretrained backbone weights (pickle-free safetensors). **`epoch_2.safetensors` is the final (3rd-epoch) checkpoint.** No FMSP adapter β€” this is the plain backbone. ### Load and generate (local clone) ```python import torch from safetensors.torch import load_file from src.model_v88_transformer import MicroMixerV88Transformer, v88_transformer_50k from src.tokenizer import ByteTokenizer # Clone the code repository first: # git clone https://github.com/llaa33219/MicroMixer-4.git && cd MicroMixer-4 cfg = v88_transformer_50k() model = MicroMixerV88Transformer(cfg) # plain backbone β€” NO attach_adapter (pretrain-only) model.load_state_dict(load_file("epoch_2.safetensors"), strict=True) model.eval() tok = ByteTokenizer() prompt = "Once upon a time, there was a little girl named Lily." ids = tok.encode(prompt) if ids and ids[-1] == tok.eos_token_id: ids = ids[:-1] # ByteTokenizer appends EOS; the prompt must end open ids = torch.tensor([ids]) with torch.no_grad(): out = model.generate( ids, max_new_tokens=200, temperature=0.0, # greedy repetition_penalty=1.2, no_repeat_ngram_size=4, eos_token_id=tok.eos_token_id, ) print(prompt + tok.decode(out[0].tolist()[len(ids):])) ``` > Note the differences from the FMSP cards: (1) **no `attach_adapter`** β€” the backbone is loaded > as-is; (2) the prompt is **raw story text**, not the `User: …\n\nAssistant: ` dialogue format. ### Load from Hugging Face Hub (no clone of the weights needed) ```python import torch from huggingface_hub import hf_hub_download from safetensors.torch import load_file from src.model_v88_transformer import MicroMixerV88Transformer, v88_transformer_50k REPO = "llaa33219/MicroT-test1-50K-TinyStories" cfg = v88_transformer_50k() model = MicroMixerV88Transformer(cfg) model.load_state_dict( load_file(hf_hub_download(REPO, "epoch_2.safetensors")), strict=True) model.eval() # ... continue a story as above ``` --- ## ⚠️ Limitations
| Limitation | Description | |------------|-------------| | **Pretrain-only β€” no instruction/QA ability** | Not FMSP-fine-tuned; it only continues TinyStories-style prose. It cannot answer questions or follow instructions. | | **Micro parameters** | 49,888 parameters; capacity is the binding constraint | | **Knows only TinyStories** | Distribution is synthetic children's stories; no real-world knowledge | | **Byte-level noise** | 256-vocab byte tokenizer; PPL not comparable to BPE baselines | | **Research use only** | Architecture/pretraining research artifact, not a production model |
--- ## 🧬 Context This is the **50K** TinyStories-pretrained arm of the **dataset-efficiency comparison study** in the [MicroMixer-4](https://github.com/llaa33219/MicroMixer-4) project β€” the pretrain-only third corpus alongside the UltraChat and SmolTalk2 FMSP arms (TinyStories is excluded from the FMSP/eval battery because it has no `User:/Assistant:` markers). Sibling repos: `llaa33219/{MicroMixer-4,MicroT-test1}-{1M..10K}-TinyStories`, plus the UltraChat/SmolTalk2 arms `…-{UltraChat,SmolTalk2}` and the discord-pretrained baselines `llaa33219/{MicroMixer-4,MicroT-test1}-{1M..10K}`. Full analysis: [DATASET_COMPARISON_ANALYSIS.md](https://github.com/llaa33219/MicroMixer-4/blob/main/DATASET_COMPARISON_ANALYSIS.md). ---
[![GitHub](https://img.shields.io/badge/Back_to_Repository-MicroMixer--4-blue?style=for-the-badge&logo=github&color=%23007BFF)](https://github.com/llaa33219/MicroMixer-4) Part of the MicroMixer-4 research project β€” V88 transformer reference (MicroT-test1), 50K preset, TinyStories pretraining (pretrain-only)