llaa33219 commited on
Commit
ee40c9b
·
verified ·
1 Parent(s): fe9394f

Add MicroT-test1-10K-UltraChat: FMSP model card + 10 epoch safetensors

Browse files
README.md ADDED
@@ -0,0 +1,327 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: pytorch
4
+ language:
5
+ - en
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - transformer
9
+ - byte-level
10
+ - causal-lm
11
+ - attention
12
+ - rope
13
+ - decoder-only
14
+ - fmsp
15
+ - micro-language-model
16
+ - sub-1m-parameters
17
+ model_name: MicroT-test1-10K-UltraChat
18
+ datasets:
19
+ - HuggingFaceH4/ultrachat_200k
20
+ - llaa33219/small-qa-en-10k
21
+ metrics:
22
+ - perplexity
23
+ - exact-match
24
+ ---
25
+
26
+ <div align="center">
27
+
28
+ <img src="https://raw.githubusercontent.com/llaa33219/MicroMixer-4/main/logo.svg" width="300" alt="MicroMixer-4 Logo"/>
29
+
30
+ # MicroT-test1-10K-UltraChat
31
+
32
+ <img src="https://img.shields.io/badge/Parameters-9%2C808-blue?style=for-the-badge&logo=python&logoColor=white&color=%23007BFF" alt="Parameters"/>
33
+ <img src="https://img.shields.io/badge/Architecture-Transformer-orange?style=for-the-badge&color=%23FF6600" alt="Architecture"/>
34
+ <img src="https://img.shields.io/badge/Fine--tuning-FMSP-green?style=for-the-badge&color=%2300D620" alt="FMSP"/>
35
+
36
+ <br/>
37
+ <br/>
38
+
39
+ <table>
40
+ <tr>
41
+ <td align="center" style="padding: 20px;">
42
+ <strong style="color: #FF6600; font-size: 1.2em;">Micro Transformer — Reference Baseline</strong><br/><em>Vanilla Attention • RoPE • Byte-Level • Decoder-Only</em>
43
+ </td>
44
+ </tr>
45
+ </table>
46
+
47
+ [![GitHub](https://img.shields.io/badge/GitHub-MicroMixer--4-blue?style=for-the-badge&logo=github&color=%23007BFF)](https://github.com/llaa33219/MicroMixer-4)
48
+
49
+ </div>
50
+
51
+ <div style="background: linear-gradient(135deg, #FF660022, #007BFF22); padding: 20px; border-radius: 10px; border-left: 4px solid #FF6600;">
52
+
53
+ ## 📋 Overview
54
+
55
+ **MicroT-test1-10K-UltraChat** is a **9,808-parameter** vanilla decoder-only **transformer** — multi-head causal self-attention with RoPE — pretrained on **UltraChat 200k** conversation data (instead of the project's Discord-Dialogues baseline) and then fine-tuned with **FMSP** on 9,012 general-knowledge QA pairs.
56
+
57
+ This is the **10K** member of the **MicroT-test1** family: the registered attention-based **reference baseline** of the MicroMixer-4 project, here rerun on UltraChat 200k as part of the **dataset-efficiency comparison study** — six parameter budgets × two architectures × two open pretraining corpora, all fine-tuned with the identical P05 FMSP recipe at seed 42. [Analysis](https://github.com/llaa33219/MicroMixer-4/blob/main/DATASET_COMPARISON_ANALYSIS.md).
58
+
59
+ It is deliberately boring — the standard 2018–2020 transformer recipe parameterized down to the sub-1M regime: no flash attention, no SwiGLU, no ALiBi, no QKNorm, no sliding window, no MQA/GQA, no MoE.
60
+
61
+ </div>
62
+
63
+ ## 🏗️ Architecture
64
+
65
+ <div align="center">
66
+
67
+ ```mermaid
68
+ graph TD
69
+ A[Byte Input] --> B[Embed 256→16]
70
+ B --> C[Transformer Block × 2]
71
+ C --> D[RMSNorm]
72
+ D --> E[LM Head Tied with Embed]
73
+ E --> F[Byte Output]
74
+
75
+ subgraph "Transformer Block (pre-norm)"
76
+ X[Input 16] --> N1[RMSNorm]
77
+ N1 --> AT["MHA 1 heads × d_head 16<br/>RoPE θ=10000 on q,k · causal SDPA"]
78
+ AT --> R1[+ residual]
79
+ R1 --> N2[RMSNorm]
80
+ N2 --> MLP["GELU MLP 16→56→16"]
81
+ MLP --> R2[+ residual]
82
+ end
83
+
84
+ style A fill:#007BFF,color:#fff
85
+ style F fill:#00D620,color:#fff
86
+ style AT fill:#FF6600,color:#fff
87
+ ```
88
+
89
+ </div>
90
+
91
+ ### Model Configuration
92
+
93
+ <table>
94
+ <tr>
95
+ <th style="background-color: #FF6600; color: white;">Parameter</th>
96
+ <th style="background-color: #007BFF; color: white;">Value</th>
97
+ </tr>
98
+ <tr><td>Hidden Dimension (d_model)</td><td><code>16</code></td></tr>
99
+ <tr><td>Attention Heads</td><td><code>1</code> (d_head = <code>16</code> at every size)</td></tr>
100
+ <tr><td>Number of Blocks</td><td><code>2</code></td></tr>
101
+ <tr><td>FFN Hidden</td><td><code>56</code></td></tr>
102
+ <tr><td>Position Encoding</td><td>RoPE θ=10000 on q/k only (non-persistent buffers)</td></tr>
103
+ <tr><td>Attention</td><td>Causal MHA via <code>F.scaled_dot_product_attention(is_causal=True)</code></td></tr>
104
+ <tr><td>Activation</td><td><code>GELU</code></td></tr>
105
+ <tr><td>Biases</td><td><b>None</b> — no bias parameters anywhere</td></tr>
106
+ <tr><td>Normalization</td><td><code>RMSNorm</code> (pre-norm)</td></tr>
107
+ <tr><td>Max Sequence Length</td><td><code>1024</code></td></tr>
108
+ <tr><td>Vocabulary Size</td><td><code>256</code> (byte-level)</td></tr>
109
+ <tr><td>Output Head</td><td>Tied with input embedding</td></tr>
110
+ </table>
111
+
112
+ ### Core Components
113
+
114
+ ```
115
+ ┌──────────────────────────────────────────────┐
116
+ │ Transformer Block (×2) │
117
+ │ h = h + MHA(RMSNorm(h)) # RoPE q/k, causal│
118
+ │ h = h + MLP(RMSNorm(h)) # GELU d→ffn→d │
119
+ │ no biases, no flash, no tricks — vanilla │
120
+ └──────────────────────────────────────────────┘
121
+ ```
122
+
123
+ The d_head=16 contract is hard-asserted across all six sizes so that attention-head behavior is comparable at every budget and never confounds the memorization measurements.
124
+
125
+ ---
126
+
127
+ ## 🎯 Generation Examples
128
+
129
+ <div style="background-color: #FF050515; padding: 15px; border-radius: 8px; border-left: 4px solid #FF6600;">
130
+
131
+ **Questions the model was trained on** (FMSP train set, 9,012 QA pairs — greedy decoding, `repetition_penalty=1.2`, `no_repeat_ngram_size=4`):
132
+
133
+ ```
134
+ [Prompt] User: Who painted the Mona Lisa?
135
+ Assistant:
136
+ [Output] The stand is is the and the the the the the the the the the the the the the the the the the the…
137
+ ```
138
+
139
+ <sub>degenerate — collapses into a repeating token loop; no real answer</sub>
140
+
141
+ ```
142
+ [Prompt] User: Who painted The Starry Night?
143
+ Assistant:
144
+ [Output] The stand is is the and the and the the the the the the the the the the the the the the the the…
145
+ ```
146
+
147
+ <sub>degenerate — collapses into a repeating token loop; no real answer</sub>
148
+
149
+
150
+ **Questions the model has never seen and cannot answer** (unanswerable probe — the correct behavior is to decline; the model's actual behavior is shown):
151
+
152
+ ```
153
+ [Prompt] User: Who painted the Glimmering Frostberry?
154
+ Assistant:
155
+ [Output] The stand a th the the the the the the the the the the the the the the the the the the the the …
156
+ ```
157
+
158
+ <sub>degenerate — collapses into a repeating token loop on out-of-distribution input</sub>
159
+
160
+ ```
161
+ [Prompt] User: Who composed the Symphony of Hollow Dawn?
162
+ Assistant:
163
+ [Output] The stand is is the and the the the the the the the the the the the the the the the the the the…
164
+ ```
165
+
166
+ <sub>degenerate — collapses into a repeating token loop on out-of-distribution input</sub>
167
+
168
+
169
+ </div>
170
+
171
+ ---
172
+
173
+ ## 📊 Results
174
+
175
+ <div style="background-color: #FF660015; padding: 15px; border-radius: 8px; border-left: 4px solid #FF6600;">
176
+
177
+ ### Pretraining (UltraChat 200k, V76 recipe, 3 epochs)
178
+
179
+ | Metric | 1 ep | 2 ep | 3 ep |
180
+ |--------|------|------|------|
181
+ | Val PPL | 7.05 | 6.80 | **6.53** |
182
+
183
+ Identical recipe to the MicroMixer-4 mixer: AdamW lr 3e-3 · WSD (warmup 500) · wd 0.01 · bs 16 · seq 1024 · seed 42.
184
+
185
+ ### FMSP fine-tuning (small-qa-en-10k, P05 recipe, 10 epochs)
186
+
187
+ | Metric | Value |
188
+ |--------|-------|
189
+ | Train QA pairs | 9,012 |
190
+ | Held-out QA pairs | 988 |
191
+ | Best-val checkpoint | `fmsp_epoch_9.safetensors` (val loss **1.7229**) |
192
+ | freeze_fraction | 0.05 (true freeze) |
193
+ | Loss | answer-only CE + probe KL (weight 0.5) |
194
+
195
+ ### Evaluation battery (post-FMSP)
196
+
197
+ | Axis | MicroT-test1-10K-UltraChat |
198
+ |------|--------|
199
+ | Chatter fluency d2 (cycles) | 0.331 (9/9) |
200
+ | Full-988 EM (seed 42) | 0 |
201
+ | Q-relevance echo / hijack % | 0.0 / 0.0 |
202
+ | OOD hijack % | 0.0% |
203
+ | Unanswerable fabrication /18 | 0 |
204
+ | Discord PPL | 9.75 |
205
+
206
+ <sub>Single-seed run (seed 42); the discord-pretrained cards report a 3-seed mean for Full-988 EM. ‡ where marked: degenerate-pass.</sub>
207
+
208
+ ### MicroT-test1 UltraChat family (same protocol, all sizes)
209
+
210
+ | Size | Params | 3ep Val PPL | Chatter d2 | Full-988 EM | qrel echo/hijack | OOD hijack |
211
+ |------|--------|-------------|------------|-------------|------------------|------------|
212
+ | 1M | 996,736 | 2.40 | 0.698 | 749 | 100.0 / 0.0 | 39.0% |
213
+ | 500K | 498,528 | 2.65 | 0.791 | 550 | 84.0 / 14.0 | 59.3% |
214
+ | 300K | 297,680 | 2.84 | 0.633 | 211 | 62.0 / 24.0 | 42.4% |
215
+ | 100K | 97,872 | 3.49 | 0.751 | 0 | 12.0 / 70.0 | 66.1% |
216
+ | 50K | 49,888 | 3.98 | 0.528 | 0 | 4.0 / 13.0 | 1.7% |
217
+ | **10K** | 9,808 | 6.53 | 0.331 | 0 | 0.0 / 0.0 | 0.0% |
218
+
219
+ <sub>Seed-42 single runs (pretrained on UltraChat 200k; the discord-pretrained families report 3-seed EM means).</sub>
220
+
221
+ </div>
222
+
223
+ ---
224
+
225
+ ## 📚 Training Data
226
+
227
+ <div style="background-color: #00D62015; padding: 15px; border-radius: 8px; border-left: 4px solid #00D620;">
228
+
229
+ 1. **Pretraining**: [UltraChat 200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k) — 146K multi-turn conversations (`train_sft` split), flattened to `User:/Assistant:` format, 1024-byte sequences, 3 epochs.
230
+ 2. **FMSP fine-tuning**: [small-qa-en-10k](https://huggingface.co/datasets/llaa33219/small-qa-en-10k) — 10K general-knowledge QA pairs (arts, science, history, geography, music…), split 9,012 train / 988 held-out. 10 epochs under the P05 recipe (5% of parameters frozen-true, answer-only CE, probe-KL 0.5).
231
+
232
+ </div>
233
+
234
+ ---
235
+
236
+ ## 🔧 Usage
237
+
238
+ ### Files in this repository
239
+ - `fmsp_epoch_{0..9}.safetensors` — per-epoch FMSP weights (pickle-free safetensors). **`fmsp_epoch_9.safetensors` is the best-val checkpoint** for this size.
240
+
241
+ ### Load and generate (local clone)
242
+
243
+ ```python
244
+ import torch
245
+ from safetensors.torch import load_file
246
+ from src.model_v88_transformer import MicroMixerV88Transformer, v88_transformer_10k
247
+ from src.fmsp import attach_adapter
248
+ from src.tokenizer import ByteTokenizer
249
+
250
+ # Clone the code repository first:
251
+ # git clone https://github.com/llaa33219/MicroMixer-4.git && cd MicroMixer-4
252
+
253
+ cfg = v88_transformer_10k()
254
+ model = MicroMixerV88Transformer(cfg)
255
+ attach_adapter(model, d_model=cfg.d_model, rank=16) # FMSP adapter (trained weights are in the file)
256
+ model.load_state_dict(load_file("fmsp_epoch_9.safetensors"), strict=True)
257
+ model.eval()
258
+
259
+ tok = ByteTokenizer()
260
+ prompt = "User: Who painted the Mona Lisa?\n\nAssistant: "
261
+ ids = tok.encode(prompt)
262
+ if ids and ids[-1] == tok.eos_token_id:
263
+ ids = ids[:-1] # ByteTokenizer appends EOS; the prompt must end open
264
+ ids = torch.tensor([ids])
265
+ with torch.no_grad():
266
+ out = model.generate(
267
+ ids, max_new_tokens=200,
268
+ temperature=0.0, # greedy — used for all reported numbers
269
+ repetition_penalty=1.2,
270
+ no_repeat_ngram_size=4,
271
+ eos_token_id=tok.eos_token_id,
272
+ )
273
+ print(tok.decode(out[0].tolist()))
274
+ ```
275
+
276
+ ### Load from Hugging Face Hub (no clone of the weights needed)
277
+
278
+ ```python
279
+ import torch
280
+ from huggingface_hub import hf_hub_download
281
+ from safetensors.torch import load_file
282
+ from src.model_v88_transformer import MicroMixerV88Transformer, v88_transformer_10k
283
+ from src.fmsp import attach_adapter
284
+
285
+ REPO = "llaa33219/MicroT-test1-10K-UltraChat"
286
+
287
+ cfg = v88_transformer_10k()
288
+ model = MicroMixerV88Transformer(cfg)
289
+ attach_adapter(model, d_model=cfg.d_model, rank=16)
290
+ model.load_state_dict(
291
+ load_file(hf_hub_download(REPO, "fmsp_epoch_9.safetensors")), strict=True)
292
+ model.eval()
293
+ # ... generate as above
294
+ ```
295
+
296
+ ---
297
+
298
+ ## ⚠️ Limitations
299
+
300
+ <div style="background-color: #FF050515; padding: 15px; border-radius: 8px; border-left: 4px solid #FF0505;">
301
+
302
+ | Limitation | Description |
303
+ |------------|-------------|
304
+ | **Micro parameters** | 9,808 parameters; capacity is the binding constraint on every axis |
305
+ | **Reference baseline, not a product** | Exists to score the Mixer against attention at matched budget |
306
+ | **Knows only what it memorized** | Knowledge is limited to the 9,012 trained QA pairs + the pretraining-corpus distribution |
307
+ | **Does not abstain** | Unknown questions are answered with fabrication or degeneration, not refusal — see the examples above |
308
+ | **Byte-level noise** | 256-vocab byte tokenizer; PPL not comparable to BPE baselines |
309
+ | **Research use only** | Architecture/scaling research artifact, not a production model |
310
+
311
+ </div>
312
+
313
+ ---
314
+
315
+ ## 🧬 Context
316
+
317
+ This is the **10K** UltraChat-pretrained arm of the **dataset-efficiency comparison study** (24 arms: {V87, V88} × {UltraChat, SmolTalk2} × six sizes) in the [MicroMixer-4](https://github.com/llaa33219/MicroMixer-4) project. Each arm swaps the Discord-Dialogues pretraining baseline for an open corpus (UltraChat 200k) and reruns the identical P05 FMSP recipe at seed 42. Sibling repos: `llaa33219/MicroMixer-4-{1M..10K}-{UltraChat,SmolTalk2}` and `llaa33219/MicroT-test1-{1M..10K}-{UltraChat,SmolTalk2}`; the discord-pretrained baselines are `llaa33219/MicroMixer-4-{1M..10K}` and `llaa33219/MicroT-test1-{1M..10K}`. Full analysis: [DATASET_COMPARISON_ANALYSIS.md](https://github.com/llaa33219/MicroMixer-4/blob/main/DATASET_COMPARISON_ANALYSIS.md).
318
+
319
+ ---
320
+
321
+ <div align="center">
322
+
323
+ [![GitHub](https://img.shields.io/badge/Back_to_Repository-MicroMixer--4-blue?style=for-the-badge&logo=github&color=%23007BFF)](https://github.com/llaa33219/MicroMixer-4)
324
+
325
+ <sub>Part of the <a href="https://github.com/llaa33219/MicroMixer-4">MicroMixer-4</a> research project — V88 transformer reference (MicroT-test1), 10K preset, UltraChat pretraining</sub>
326
+
327
+ </div>
fmsp_epoch_0.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d95acd964c51c9a4e5aef98a761c7efa0145191362ba3acb2a149b3dfacf6824
3
+ size 43008
fmsp_epoch_1.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:632d7d87e6b22305dd32bb22e0c1f835b2e5d9159f447f7ec98e0db3f894d2dc
3
+ size 43008
fmsp_epoch_2.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c0cb8742b2598847a9993637d2937b24b7692365991e25bf87e4df0964a6bb93
3
+ size 43008
fmsp_epoch_3.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d16ae84ad821a8dbe3b3ff1a34380995c5a6fc235ed40c57000d17b2bad7994d
3
+ size 43008
fmsp_epoch_4.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:39e2ebf4a0710d002faab651eaebde76dcbab82fd9e46dd952b5daac390e6f78
3
+ size 43008
fmsp_epoch_5.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4651c5563492ec81641003fac348e0c10c01028f0488e9ebba56d2e345e4158e
3
+ size 43008
fmsp_epoch_6.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8eb4497e0890313649307c17c719a4f71b554986bab9f8ec44d889372ab9abd5
3
+ size 43008
fmsp_epoch_7.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5dab7cc42cacf85b4210e7aeb4b9aacf9d4da244f40c32aa32efbc769bcc8c46
3
+ size 43008
fmsp_epoch_8.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7d9bd14248c15f552e85845a0c95c8b02c5f17edab7575246847bda01b95907d
3
+ size 43008
fmsp_epoch_9.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7e5cfaa43195e5373083be6caa451b284624edbd7dbfa208dad1d8ba83d4de37
3
+ size 43008