llaa33219 commited on
Commit
38f1a42
·
verified ·
1 Parent(s): d2bffce

Add MicroMixer-4-50K-TinyStories: TinyStories pretrain card + 3 epoch safetensors

Browse files
Files changed (4) hide show
  1. README.md +303 -0
  2. epoch_0.safetensors +3 -0
  3. epoch_1.safetensors +3 -0
  4. epoch_2.safetensors +3 -0
README.md ADDED
@@ -0,0 +1,303 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: pytorch
4
+ language:
5
+ - en
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - mlp-mixer
9
+ - byte-level
10
+ - causal-lm
11
+ - attention-free
12
+ - ccd-mixer
13
+ - content-gated-dilated-conv
14
+ - tiny-stories
15
+ - pretrained-backbone
16
+ - micro-language-model
17
+ - sub-1m-parameters
18
+ model_name: MicroMixer-4-50K-TinyStories
19
+ datasets:
20
+ - roneneldan/TinyStories
21
+ metrics:
22
+ - perplexity
23
+ ---
24
+
25
+ <div align="center">
26
+
27
+ <img src="https://raw.githubusercontent.com/llaa33219/MicroMixer-4/main/logo.svg" width="300" alt="MicroMixer-4 Logo"/>
28
+
29
+ # MicroMixer-4-50K-TinyStories
30
+
31
+ <img src="https://img.shields.io/badge/Parameters-48%2C684-blue?style=for-the-badge&logo=python&logoColor=white&color=%23007BFF" alt="Parameters"/>
32
+ <img src="https://img.shields.io/badge/Architecture-CCD--Mixer-purple?style=for-the-badge&color=%23AE00FF" alt="Architecture"/>
33
+ <img src="https://img.shields.io/badge/Stage-Pretrain-green?style=for-the-badge&color=%2300D620" alt="Pretrain"/>
34
+
35
+ <br/>
36
+ <br/>
37
+
38
+ <table>
39
+ <tr>
40
+ <td align="center" style="padding: 20px;">
41
+ <strong style="color: #007BFF; font-size: 1.2em;">Micro Language Model</strong><br/><em>Attention-Free • MLP-Only • Byte-Level • Content-Gated Dilated Convolution</em>
42
+ </td>
43
+ </tr>
44
+ </table>
45
+
46
+ [![GitHub](https://img.shields.io/badge/GitHub-MicroMixer--4-blue?style=for-the-badge&logo=github&color=%23007BFF)](https://github.com/llaa33219/MicroMixer-4)
47
+
48
+ </div>
49
+
50
+ <div style="background: linear-gradient(135deg, #007BFF22, #AE00FF22); padding: 20px; border-radius: 10px; border-left: 4px solid #007BFF;">
51
+
52
+ ## 📋 Overview
53
+
54
+ **MicroMixer-4-50K-TinyStories** is a **48,684-parameter** pure MLP-Mixer causal language model — **no attention, no recurrence, no SSM** — **pretrained on [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories)** as part of the MicroMixer-4 **dataset-efficiency comparison study** ([analysis](https://github.com/llaa33219/MicroMixer-4/blob/main/DATASET_COMPARISON_ANALYSIS.md)). TinyStories is a corpus of ~200K short synthetic children's stories (GPT-3.5/4-generated, constrained grammar and vocabulary).
55
+
56
+ ⚠️ This repo is **pretrain-only** — it is **NOT FMSP-fine-tuned**. TinyStories contains free-flowing prose with no `User:`/`Assistant:` dialogue markers, so the FMSP answer-only cross-entropy recipe does not apply (there is no answer region to isolate). What you get is the **raw pretrained backbone** (seed 42, V76 recipe, 3 epochs). It continues stories; it does not answer questions or follow instructions.
57
+
58
+ The backbone is **V87 Final**, the project's champion CCD-Mixer architecture (the V83-RPG champion frozen and scaled to six budgets). The 50K preset reproduces the champion recipe at its budget.
59
+
60
+ </div>
61
+
62
+ ## 🏗️ Architecture
63
+
64
+ <div align="center">
65
+
66
+ ```mermaid
67
+ graph TD
68
+ A[Byte Input] --> B[Embed 256→32 NoPE]
69
+ B --> C[CCD-Mixer Block × 4]
70
+ C --> D[RMSNorm]
71
+ D --> E[LM Head Tied with Embed]
72
+ E --> F[Byte Output]
73
+
74
+ subgraph "CCD-Mixer Block"
75
+ X[Input 32] --> U["Linear d→2d → split v, g"]
76
+ U --> RP[Full RoPE on v AND g]
77
+ RP --> M["Shared-weight dilated conv<br/>dilations 1·2·4·8, k=65"]
78
+ M --> G["Per-position 4-way gate<br/>softmax(Linear_dil(x)/τ)"]
79
+ G --> O["W_o(v ⊙ g) — zero-init"]
80
+ O --> SW[SwiGLU Channel-Mix]
81
+ SW --> RM[ReMixerLayer sidecar]
82
+ end
83
+
84
+ style A fill:#007BFF,color:#fff
85
+ style F fill:#00D620,color:#fff
86
+ style G fill:#AE00FF,color:#fff
87
+ style M fill:#FF6600,color:#fff
88
+ ```
89
+
90
+ </div>
91
+
92
+ ### Model Configuration
93
+
94
+ <table>
95
+ <tr>
96
+ <th style="background-color: #007BFF; color: white;">Parameter</th>
97
+ <th style="background-color: #AE00FF; color: white;">Value</th>
98
+ </tr>
99
+ <tr><td>Hidden Dimension (d_model)</td><td><code>32</code></td></tr>
100
+ <tr><td>Number of Blocks</td><td><code>4</code></td></tr>
101
+ <tr><td>Token-Mix</td><td><code>GLCTokenMixCCD</code> (content-gated mixture of shared-weight dilated causal conv)</td></tr>
102
+ <tr><td>Dilations</td><td><code>(1, 2, 4, 8)</code> — one shared depthwise kernel, zero extra conv params</td></tr>
103
+ <tr><td>Depthwise Kernel Size</td><td><code>65</code></td></tr>
104
+ <tr><td>RoPE</td><td>Full RoPE on <b>both</b> v and g (V76 "RPG" pattern)</td></tr>
105
+ <tr><td>Channel-Mix</td><td><code>SwiGLU</code></td></tr>
106
+ <tr><td>Sidecar</td><td><code>ReMixerLayer</code> per block (label_dim 16, pool_heads 4)</td></tr>
107
+ <tr><td>Max Sequence Length</td><td><code>1024</code></td></tr>
108
+ <tr><td>Vocabulary Size</td><td><code>256</code> (byte-level)</td></tr>
109
+ <tr><td>Position Encoding</td><td>RoPE inside token-mix only; no position embedding table</td></tr>
110
+ <tr><td>Normalization</td><td><code>RMSNorm</code> (pre-norm)</td></tr>
111
+ <tr><td>Output Head</td><td>Tied with input embedding</td></tr>
112
+ <tr><td>Zero-Init</td><td><code>W_o</code>, <code>dil_gate</code>, <code>log_τ</code> — silent at init</td></tr>
113
+ </table>
114
+
115
+ ### Core Components
116
+
117
+ ```
118
+ ┌─────────────────────────────────��────────────────────────────┐
119
+ │ CCD-Mixer Block (×4) │
120
+ │ u = Linear(d → 2d)(x) │
121
+ │ v, g = u.chunk(2) │
122
+ │ v = RoPE(v) g = RoPE(g) ← full-RoPE (RPG) │
123
+ │ y_d = CausalDSConv(v, dilation=d) for d ∈ (1,2,4,8) │
124
+ │ └── ONE shared depthwise kernel │
125
+ │ w(t) = softmax(Linear_dil(x)_t / τ) ← per-position │
126
+ │ v = Σ_d w_d(t) · y_d(t) time-varying filter │
127
+ │ out = W_o(v ⊙ g) ← W_o zero-init │
128
+ │ then SwiGLU channel-mix + ReMixerLayer sidecar │
129
+ └──────────────────────────────────────────────────────────────┘
130
+ ```
131
+
132
+ The token-mix is **non-LTI** (time-varying): the per-position gate remixes four dilated views of the same kernel at every byte, which is the mechanism that breaks the periodic-orbit collapse that pure LTI mixers fall into — without attention and without a position table.
133
+
134
+ ---
135
+
136
+ ## 🎯 Generation Examples
137
+
138
+ <div style="background-color: #FF050515; padding: 15px; border-radius: 8px; border-left: 4px solid #FF6600;">
139
+
140
+ **Story-continuation probes** (raw text — no `User:/Assistant:` prefix, since TinyStories is a plain LM; greedy decoding, `repetition_penalty=1.2`, `no_repeat_ngram_size=4`, `max_new_tokens=200`):
141
+
142
+ ```
143
+ [Prompt] Once upon a time, there was a little girl named Lily.
144
+ [Output] She loved to play with her friends. One day, Lily was very happy and saw a big box of the store…
145
+ ```
146
+
147
+ <sub>coherent — on-topic story continuation in the TinyStories register</sub>
148
+
149
+ ```
150
+ [Prompt] One day, a boy named Tom found a big red balloon.
151
+ [Output] They were very happy and he was so happy. One day, he was so happy to help him and said, "I wan…
152
+ ```
153
+
154
+ <sub>coherent — on-topic story continuation in the TinyStories register</sub>
155
+
156
+ ```
157
+ [Prompt] The sun was shining and the birds were singing.
158
+ [Output] The bird was so happy and said, "I want to see the store with the store with the store with the…
159
+ ```
160
+
161
+ <sub>drifts — starts as story but falls into a repetitive loop</sub>
162
+
163
+ ```
164
+ [Prompt] A little dog ran into the garden and saw
165
+ [Output] a big box of fun. He was so happy and he wanted to see the box and said, "I want to see the bea…
166
+ ```
167
+
168
+ <sub>coherent — on-topic story continuation in the TinyStories register</sub>
169
+
170
+
171
+ </div>
172
+
173
+ ---
174
+
175
+ ## 📊 Results
176
+
177
+ <div style="background-color: #007BFF15; padding: 15px; border-radius: 8px; border-left: 4px solid #007BFF;">
178
+
179
+ ### Pretraining (TinyStories, V76 recipe, 3 epochs)
180
+
181
+ | Metric | 1 ep | 2 ep | 3 ep |
182
+ |--------|------|------|------|
183
+ | Val PPL | 2.83 | 2.75 | **2.60** |
184
+
185
+ AdamW lr 3e-3 · WSD (warmup 500) · wd 0.01 · bs 16 · seq 1024 · seed 42 · plain CE on non-pad bytes.
186
+
187
+ ### MicroMixer-4 TinyStories family (pretrain-only, all sizes)
188
+
189
+ | Size | Params | 3ep Val PPL |
190
+ |------|--------|------------|
191
+ | 1M | 996,873 | **1.78** |
192
+ | 500K | 491,742 | **1.90** |
193
+ | 300K | 292,525 | **2.00** |
194
+ | 100K | 95,084 | **2.31** |
195
+ | **50K** | 48,684 | **2.60** |
196
+ | 10K | 9,666 | **4.06** |
197
+
198
+ <sub>Pretrain-only family — FMSP-based axes (chatter fluency, full-988 EM, q-relevance, OOD, unanswerable fabrication) are N/A: TinyStories has no `User:/Assistant:` markers, so the answer-only-CE recipe does not apply.</sub>
199
+
200
+ </div>
201
+
202
+ ---
203
+
204
+ ## 📚 Training Data
205
+
206
+ <div style="background-color: #00D62015; padding: 15px; border-radius: 8px; border-left: 4px solid #00D620;">
207
+
208
+ 1. **Pretraining**: [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories) — synthetic short stories generated by GPT-3.5/4 with a constrained vocabulary and simple grammar, ~200K stories sampled, flattened to 1024-byte sequences, 3 epochs. No `User:/Assistant:` dialogue structure.
209
+
210
+ </div>
211
+
212
+ ---
213
+
214
+ ## 🔧 Usage
215
+
216
+ ### Files in this repository
217
+ - `epoch_{0,1,2}.safetensors` — per-epoch pretrained backbone weights (pickle-free safetensors). **`epoch_2.safetensors` is the final (3rd-epoch) checkpoint.** No FMSP adapter — this is the plain backbone.
218
+
219
+ ### Load and generate (local clone)
220
+
221
+ ```python
222
+ import torch
223
+ from safetensors.torch import load_file
224
+ from src.model_v87_final import MicroMixerV87Final, v87_final_50k
225
+ from src.tokenizer import ByteTokenizer
226
+
227
+ # Clone the code repository first:
228
+ # git clone https://github.com/llaa33219/MicroMixer-4.git && cd MicroMixer-4
229
+
230
+ cfg = v87_final_50k()
231
+ model = MicroMixerV87Final(cfg) # plain backbone — NO attach_adapter (pretrain-only)
232
+ model.load_state_dict(load_file("epoch_2.safetensors"), strict=True)
233
+ model.eval()
234
+
235
+ tok = ByteTokenizer()
236
+ prompt = "Once upon a time, there was a little girl named Lily."
237
+ ids = tok.encode(prompt)
238
+ if ids and ids[-1] == tok.eos_token_id:
239
+ ids = ids[:-1] # ByteTokenizer appends EOS; the prompt must end open
240
+ ids = torch.tensor([ids])
241
+ with torch.no_grad():
242
+ out = model.generate(
243
+ ids, max_new_tokens=200,
244
+ temperature=0.0, # greedy
245
+ repetition_penalty=1.2,
246
+ no_repeat_ngram_size=4,
247
+ eos_token_id=tok.eos_token_id,
248
+ )
249
+ print(prompt + tok.decode(out[0].tolist()[len(ids):]))
250
+ ```
251
+
252
+ > Note the differences from the FMSP cards: (1) **no `attach_adapter`** — the backbone is loaded
253
+ > as-is; (2) the prompt is **raw story text**, not the `User: …\n\nAssistant: ` dialogue format.
254
+
255
+ ### Load from Hugging Face Hub (no clone of the weights needed)
256
+
257
+ ```python
258
+ import torch
259
+ from huggingface_hub import hf_hub_download
260
+ from safetensors.torch import load_file
261
+ from src.model_v87_final import MicroMixerV87Final, v87_final_50k
262
+
263
+ REPO = "llaa33219/MicroMixer-4-50K-TinyStories"
264
+
265
+ cfg = v87_final_50k()
266
+ model = MicroMixerV87Final(cfg)
267
+ model.load_state_dict(
268
+ load_file(hf_hub_download(REPO, "epoch_2.safetensors")), strict=True)
269
+ model.eval()
270
+ # ... continue a story as above
271
+ ```
272
+
273
+ ---
274
+
275
+ ## ⚠️ Limitations
276
+
277
+ <div style="background-color: #FF050515; padding: 15px; border-radius: 8px; border-left: 4px solid #FF0505;">
278
+
279
+ | Limitation | Description |
280
+ |------------|-------------|
281
+ | **Pretrain-only — no instruction/QA ability** | Not FMSP-fine-tuned; it only continues TinyStories-style prose. It cannot answer questions or follow instructions. |
282
+ | **Micro parameters** | 48,684 parameters; capacity is the binding constraint |
283
+ | **Knows only TinyStories** | Distribution is synthetic children's stories; no real-world knowledge |
284
+ | **Byte-level noise** | 256-vocab byte tokenizer; PPL not comparable to BPE baselines |
285
+ | **Research use only** | Architecture/pretraining research artifact, not a production model |
286
+
287
+ </div>
288
+
289
+ ---
290
+
291
+ ## 🧬 Context
292
+
293
+ This is the **50K** TinyStories-pretrained arm of the **dataset-efficiency comparison study** in the [MicroMixer-4](https://github.com/llaa33219/MicroMixer-4) project — the pretrain-only third corpus alongside the UltraChat and SmolTalk2 FMSP arms (TinyStories is excluded from the FMSP/eval battery because it has no `User:/Assistant:` markers). Sibling repos: `llaa33219/{MicroMixer-4,MicroT-test1}-{1M..10K}-TinyStories`, plus the UltraChat/SmolTalk2 arms `…-{UltraChat,SmolTalk2}` and the discord-pretrained baselines `llaa33219/{MicroMixer-4,MicroT-test1}-{1M..10K}`. Full analysis: [DATASET_COMPARISON_ANALYSIS.md](https://github.com/llaa33219/MicroMixer-4/blob/main/DATASET_COMPARISON_ANALYSIS.md).
294
+
295
+ ---
296
+
297
+ <div align="center">
298
+
299
+ [![GitHub](https://img.shields.io/badge/Back_to_Repository-MicroMixer--4-blue?style=for-the-badge&logo=github&color=%23007BFF)](https://github.com/llaa33219/MicroMixer-4)
300
+
301
+ <sub>Part of the <a href="https://github.com/llaa33219/MicroMixer-4">MicroMixer-4</a> research project — V87 Final (CCD-Mixer) family, 50K preset, TinyStories pretraining (pretrain-only)</sub>
302
+
303
+ </div>
epoch_0.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:192b1bd630a8fb7fb417b431f65751abe2700ada4484bd92d6a8b8d4705395ba
3
+ size 221480
epoch_1.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:20e26d3adaf723c62745e6cf2f60cc2faa53c96305a59ae32327d5ef601b7533
3
+ size 221480
epoch_2.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ee7ca038a77ef1bb805e419305a54e5581be2c2e0d4c64332c6850d8285f3615
3
+ size 221480