llaa33219 commited on
Commit
78e8604
·
verified ·
1 Parent(s): 3468440

Add MicroMixer-4-100K-UltraChat: FMSP model card + 10 epoch safetensors

Browse files
README.md ADDED
@@ -0,0 +1,336 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: pytorch
4
+ language:
5
+ - en
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - mlp-mixer
9
+ - byte-level
10
+ - causal-lm
11
+ - attention-free
12
+ - ccd-mixer
13
+ - content-gated-dilated-conv
14
+ - fmsp
15
+ - micro-language-model
16
+ - sub-1m-parameters
17
+ model_name: MicroMixer-4-100K-UltraChat
18
+ datasets:
19
+ - HuggingFaceH4/ultrachat_200k
20
+ - llaa33219/small-qa-en-10k
21
+ metrics:
22
+ - perplexity
23
+ - exact-match
24
+ ---
25
+
26
+ <div align="center">
27
+
28
+ <img src="https://raw.githubusercontent.com/llaa33219/MicroMixer-4/main/logo.svg" width="300" alt="MicroMixer-4 Logo"/>
29
+
30
+ # MicroMixer-4-100K-UltraChat
31
+
32
+ <img src="https://img.shields.io/badge/Parameters-95%2C084-blue?style=for-the-badge&logo=python&logoColor=white&color=%23007BFF" alt="Parameters"/>
33
+ <img src="https://img.shields.io/badge/Architecture-CCD--Mixer-purple?style=for-the-badge&color=%23AE00FF" alt="Architecture"/>
34
+ <img src="https://img.shields.io/badge/Fine--tuning-FMSP-green?style=for-the-badge&color=%2300D620" alt="FMSP"/>
35
+
36
+ <br/>
37
+ <br/>
38
+
39
+ <table>
40
+ <tr>
41
+ <td align="center" style="padding: 20px;">
42
+ <strong style="color: #007BFF; font-size: 1.2em;">Micro Language Model</strong><br/><em>Attention-Free • MLP-Only • Byte-Level • Content-Gated Dilated Convolution</em>
43
+ </td>
44
+ </tr>
45
+ </table>
46
+
47
+ [![GitHub](https://img.shields.io/badge/GitHub-MicroMixer--4-blue?style=for-the-badge&logo=github&color=%23007BFF)](https://github.com/llaa33219/MicroMixer-4)
48
+
49
+ </div>
50
+
51
+ <div style="background: linear-gradient(135deg, #007BFF22, #AE00FF22); padding: 20px; border-radius: 10px; border-left: 4px solid #007BFF;">
52
+
53
+ ## 📋 Overview
54
+
55
+ **MicroMixer-4-100K-UltraChat** is a **95,084-parameter** pure MLP-Mixer causal language model — **no attention, no recurrence, no SSM** — pretrained on **UltraChat 200k** conversation data (instead of the project's Discord-Dialogues baseline) and then fine-tuned with **FMSP** on 9,012 general-knowledge QA pairs.
56
+
57
+ This is the **100K** member of the MicroMixer-4 (V87 Final) family: part of the project's **dataset-efficiency comparison study** — six parameter budgets × two architectures × two open pretraining corpora (UltraChat 200k and SmolTalk2), all fine-tuned with the identical P05 FMSP recipe at seed 42. [Analysis](https://github.com/llaa33219/MicroMixer-4/blob/main/DATASET_COMPARISON_ANALYSIS.md).
58
+
59
+ The backbone is **V87 Final**, the project's champion architecture — a CCD-Mixer (Content-gated mixture of shared-weight Dilated convolutions) crowned overall champion of the 1M architecture census (V86), frozen as the final chassis and scaled to six parameter budgets. The 100K preset reproduces the champion recipe verbatim at its budget.
60
+
61
+ </div>
62
+
63
+ ## 🏗️ Architecture
64
+
65
+ <div align="center">
66
+
67
+ ```mermaid
68
+ graph TD
69
+ A[Byte Input] --> B[Embed 256→48 NoPE]
70
+ B --> C[CCD-Mixer Block × 4]
71
+ C --> D[RMSNorm]
72
+ D --> E[LM Head Tied with Embed]
73
+ E --> F[Byte Output]
74
+
75
+ subgraph "CCD-Mixer Block"
76
+ X[Input 48] --> U["Linear d→2d → split v, g"]
77
+ U --> RP[Full RoPE on v AND g]
78
+ RP --> M["Shared-weight dilated conv<br/>dilations 1·2·4·8, k=65"]
79
+ M --> G["Per-position 4-way gate<br/>softmax(Linear_dil(x)/τ)"]
80
+ G --> O["W_o(v ⊙ g) — zero-init"]
81
+ O --> SW[SwiGLU Channel-Mix]
82
+ SW --> RM[ReMixerLayer sidecar]
83
+ end
84
+
85
+ style A fill:#007BFF,color:#fff
86
+ style F fill:#00D620,color:#fff
87
+ style G fill:#AE00FF,color:#fff
88
+ style M fill:#FF6600,color:#fff
89
+ ```
90
+
91
+ </div>
92
+
93
+ ### Model Configuration
94
+
95
+ <table>
96
+ <tr>
97
+ <th style="background-color: #007BFF; color: white;">Parameter</th>
98
+ <th style="background-color: #AE00FF; color: white;">Value</th>
99
+ </tr>
100
+ <tr><td>Hidden Dimension (d_model)</td><td><code>48</code></td></tr>
101
+ <tr><td>Number of Blocks</td><td><code>4</code></td></tr>
102
+ <tr><td>Token-Mix</td><td><code>GLCTokenMixCCD</code> (content-gated mixture of shared-weight dilated causal conv)</td></tr>
103
+ <tr><td>Dilations</td><td><code>(1, 2, 4, 8)</code> — one shared depthwise kernel, zero extra conv params</td></tr>
104
+ <tr><td>Depthwise Kernel Size</td><td><code>65</code></td></tr>
105
+ <tr><td>RoPE</td><td>Full RoPE on <b>both</b> v and g (V76 "RPG" pattern)</td></tr>
106
+ <tr><td>Channel-Mix</td><td><code>SwiGLU</code></td></tr>
107
+ <tr><td>Sidecar</td><td><code>ReMixerLayer</code> per block (label_dim 16, pool_heads 4)</td></tr>
108
+ <tr><td>Max Sequence Length</td><td><code>1024</code></td></tr>
109
+ <tr><td>Vocabulary Size</td><td><code>256</code> (byte-level)</td></tr>
110
+ <tr><td>Position Encoding</td><td>RoPE inside token-mix only; no position embedding table</td></tr>
111
+ <tr><td>Normalization</td><td><code>RMSNorm</code> (pre-norm)</td></tr>
112
+ <tr><td>Output Head</td><td>Tied with input embedding</td></tr>
113
+ <tr><td>Zero-Init</td><td><code>W_o</code>, <code>dil_gate</code>, <code>log_τ</code> — silent at init</td></tr>
114
+ </table>
115
+
116
+ ### Core Components
117
+
118
+ ```
119
+ ┌────────────────────────────────────────────────────��─────────┐
120
+ │ CCD-Mixer Block (×4) │
121
+ │ u = Linear(d → 2d)(x) │
122
+ │ v, g = u.chunk(2) │
123
+ │ v = RoPE(v) g = RoPE(g) ← full-RoPE (RPG) │
124
+ │ y_d = CausalDSConv(v, dilation=d) for d ∈ (1,2,4,8) │
125
+ │ └── ONE shared depthwise kernel │
126
+ │ w(t) = softmax(Linear_dil(x)_t / τ) ← per-position │
127
+ │ v = Σ_d w_d(t) · y_d(t) time-varying filter │
128
+ │ out = W_o(v ⊙ g) ← W_o zero-init │
129
+ │ then SwiGLU channel-mix + ReMixerLayer sidecar │
130
+ └──────────────────────────────────────────────────────────────┘
131
+ ```
132
+
133
+ The token-mix is **non-LTI** (time-varying): the per-position gate remixes four dilated views of the same kernel at every byte, which is the mechanism that breaks the periodic-orbit collapse that pure LTI mixers fall into — without attention and without a position table.
134
+
135
+ ---
136
+
137
+ ## 🎯 Generation Examples
138
+
139
+ <div style="background-color: #FF050515; padding: 15px; border-radius: 8px; border-left: 4px solid #FF6600;">
140
+
141
+ **Questions the model was trained on** (FMSP train set, 9,012 QA pairs — greedy decoding, `repetition_penalty=1.2`, `no_repeat_ngram_size=4`):
142
+
143
+ ```
144
+ [Prompt] User: Who painted the Mona Lisa?
145
+ Assistant:
146
+ [Output] The Sun is the largest of the Sun, which common structure and the Sun.
147
+ ```
148
+
149
+ <sub>fabricates — gives a wrong answer to a trained question at this size</sub>
150
+
151
+ ```
152
+ [Prompt] User: Who painted The Starry Night?
153
+ Assistant:
154
+ [Output] The Sun is the largest of the Sun, which common strings of his falling through the Sun, which c…
155
+ ```
156
+
157
+ <sub>fabricates — gives a wrong answer to a trained question at this size</sub>
158
+
159
+
160
+ **Questions the model has never seen and cannot answer** (unanswerable probe — the correct behavior is to decline; the model's actual behavior is shown):
161
+
162
+ ```
163
+ [Prompt] User: Who painted the Glimmering Frostberry?
164
+ Assistant:
165
+ [Output] The Sun, which was a standard on the Sun, which common strings that allows the story of the wor…
166
+ ```
167
+
168
+ <sub>fabricates — plausible-sounding nonsense on a nonexistent subject</sub>
169
+
170
+ ```
171
+ [Prompt] User: Who composed the Symphony of Hollow Dawn?
172
+ Assistant:
173
+ [Output] The Sun is the largest of the Sun, which combines the first personal information and inventing …
174
+ ```
175
+
176
+ <sub>fabricates — plausible-sounding nonsense on a nonexistent subject</sub>
177
+
178
+
179
+ </div>
180
+
181
+ ---
182
+
183
+ ## 📊 Results
184
+
185
+ <div style="background-color: #007BFF15; padding: 15px; border-radius: 8px; border-left: 4px solid #007BFF;">
186
+
187
+ ### Pretraining (UltraChat 200k, V76 recipe, 3 epochs)
188
+
189
+ | Metric | 1 ep | 2 ep | 3 ep |
190
+ |--------|------|------|------|
191
+ | Val PPL | 3.85 | 3.74 | **3.51** |
192
+
193
+ AdamW lr 3e-3 · WSD (warmup 500) · wd 0.01 · bs 16 · seq 1024 · seed 42 · plain CE on non-pad bytes.
194
+
195
+ ### FMSP fine-tuning (small-qa-en-10k, P05 recipe, 10 epochs)
196
+
197
+ | Metric | Value |
198
+ |--------|-------|
199
+ | Train QA pairs | 9,012 |
200
+ | Held-out QA pairs | 988 |
201
+ | Best-val checkpoint | `fmsp_epoch_9.safetensors` (val loss **0.7425**) |
202
+ | freeze_fraction | 0.05 (true freeze) |
203
+ | Loss | answer-only CE + probe KL (weight 0.5) |
204
+
205
+ ### Evaluation battery (post-FMSP)
206
+
207
+ | Axis | MicroMixer-4-100K-UltraChat |
208
+ |------|--------|
209
+ | Chatter fluency d2 (cycles) | 0.521 (3/9) |
210
+ | Full-988 EM (seed 42) | 0 |
211
+ | Q-relevance echo / hijack % | 7.0 / 75.0 |
212
+ | OOD hijack % | 61.0% |
213
+ | Unanswerable fabrication /18 | 16 |
214
+ | Discord PPL | 11.38 |
215
+
216
+ <sub>Single-seed run (seed 42); the discord-pretrained cards report a 3-seed mean for Full-988 EM. ‡ where marked: degenerate-pass.</sub>
217
+
218
+ ### MicroMixer-4 UltraChat family (same protocol, all sizes)
219
+
220
+ | Size | Params | 3ep Val PPL | Chatter d2 | Full-988 EM | qrel echo/hijack | OOD hijack |
221
+ |------|--------|-------------|------------|-------------|------------------|------------|
222
+ | 1M | 996,873 | 2.55 | 0.842 | 694 | 99.0 / 1.0 | 64.4% |
223
+ | 500K | 491,742 | 2.75 | 0.701 | 423 | 53.0 / 39.0 | 52.5% |
224
+ | 300K | 292,525 | 2.94 | 0.868 | 134 | 16.0 / 65.0 | 66.1% |
225
+ | **100K** | 95,084 | 3.51 | 0.521 | 0 | 7.0 / 75.0 | 61.0% |
226
+ | 50K | 48,684 | 4.03 | 0.428 | 0 | 3.0 / 22.0 | 18.6% |
227
+ | 10K | 9,666 | 6.38 | 0.338 | 0 | 0.0 / 1.0 | 0.0% |
228
+
229
+ <sub>Seed-42 single runs (pretrained on UltraChat 200k; the discord-pretrained families report 3-seed EM means).</sub>
230
+
231
+ </div>
232
+
233
+ ---
234
+
235
+ ## 📚 Training Data
236
+
237
+ <div style="background-color: #00D62015; padding: 15px; border-radius: 8px; border-left: 4px solid #00D620;">
238
+
239
+ 1. **Pretraining**: [UltraChat 200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k) — 146K multi-turn conversations (`train_sft` split), flattened to `User:/Assistant:` format, 1024-byte sequences, 3 epochs.
240
+ 2. **FMSP fine-tuning**: [small-qa-en-10k](https://huggingface.co/datasets/llaa33219/small-qa-en-10k) — 10K general-knowledge QA pairs (arts, science, history, geography, music…), split 9,012 train / 988 held-out. 10 epochs under the P05 recipe (5% of parameters frozen-true, answer-only CE, probe-KL 0.5).
241
+
242
+ </div>
243
+
244
+ ---
245
+
246
+ ## 🔧 Usage
247
+
248
+ ### Files in this repository
249
+ - `fmsp_epoch_{0..9}.safetensors` — per-epoch FMSP weights (pickle-free safetensors). **`fmsp_epoch_9.safetensors` is the best-val checkpoint** for this size.
250
+
251
+ ### Load and generate (local clone)
252
+
253
+ ```python
254
+ import torch
255
+ from safetensors.torch import load_file
256
+ from src.model_v87_final import MicroMixerV87Final, v87_final_100k
257
+ from src.fmsp import attach_adapter
258
+ from src.tokenizer import ByteTokenizer
259
+
260
+ # Clone the code repository first:
261
+ # git clone https://github.com/llaa33219/MicroMixer-4.git && cd MicroMixer-4
262
+
263
+ cfg = v87_final_100k()
264
+ model = MicroMixerV87Final(cfg)
265
+ attach_adapter(model, d_model=cfg.d_model, rank=16) # FMSP adapter (trained weights are in the file)
266
+ model.load_state_dict(load_file("fmsp_epoch_9.safetensors"), strict=True)
267
+ model.eval()
268
+
269
+ tok = ByteTokenizer()
270
+ prompt = "User: Who painted the Mona Lisa?\n\nAssistant: "
271
+ ids = tok.encode(prompt)
272
+ if ids and ids[-1] == tok.eos_token_id:
273
+ ids = ids[:-1] # ByteTokenizer appends EOS; the prompt must end open
274
+ ids = torch.tensor([ids])
275
+ with torch.no_grad():
276
+ out = model.generate(
277
+ ids, max_new_tokens=200,
278
+ temperature=0.0, # greedy — used for all reported numbers
279
+ repetition_penalty=1.2,
280
+ no_repeat_ngram_size=4,
281
+ eos_token_id=tok.eos_token_id,
282
+ )
283
+ print(tok.decode(out[0].tolist()))
284
+ ```
285
+
286
+ ### Load from Hugging Face Hub (no clone of the weights needed)
287
+
288
+ ```python
289
+ import torch
290
+ from huggingface_hub import hf_hub_download
291
+ from safetensors.torch import load_file
292
+ from src.model_v87_final import MicroMixerV87Final, v87_final_100k
293
+ from src.fmsp import attach_adapter
294
+
295
+ REPO = "llaa33219/MicroMixer-4-100K-UltraChat"
296
+
297
+ cfg = v87_final_100k()
298
+ model = MicroMixerV87Final(cfg)
299
+ attach_adapter(model, d_model=cfg.d_model, rank=16)
300
+ model.load_state_dict(
301
+ load_file(hf_hub_download(REPO, "fmsp_epoch_9.safetensors")), strict=True)
302
+ model.eval()
303
+ # ... generate as above
304
+ ```
305
+
306
+ ---
307
+
308
+ ## ⚠️ Limitations
309
+
310
+ <div style="background-color: #FF050515; padding: 15px; border-radius: 8px; border-left: 4px solid #FF0505;">
311
+
312
+ | Limitation | Description |
313
+ |------------|-------------|
314
+ | **Micro parameters** | 95,084 parameters; capacity is the binding constraint on every axis |
315
+ | **Knows only what it memorized** | Knowledge is limited to the 9,012 trained QA pairs + the pretraining-corpus distribution |
316
+ | **Does not abstain** | Unknown questions are answered with fabrication or degeneration, not refusal — see the examples above |
317
+ | **Byte-level noise** | 256-vocab byte tokenizer; PPL not comparable to BPE baselines |
318
+ | **Research use only** | Architecture/scaling research artifact, not a production model |
319
+
320
+ </div>
321
+
322
+ ---
323
+
324
+ ## 🧬 Context
325
+
326
+ This is the **100K** UltraChat-pretrained arm of the **dataset-efficiency comparison study** (24 arms: {V87, V88} × {UltraChat, SmolTalk2} × six sizes) in the [MicroMixer-4](https://github.com/llaa33219/MicroMixer-4) project. Each arm swaps the Discord-Dialogues pretraining baseline for an open corpus (UltraChat 200k) and reruns the identical P05 FMSP recipe at seed 42. Sibling repos: `llaa33219/MicroMixer-4-{1M..10K}-{UltraChat,SmolTalk2}` and `llaa33219/MicroT-test1-{1M..10K}-{UltraChat,SmolTalk2}`; the discord-pretrained baselines are `llaa33219/MicroMixer-4-{1M..10K}` and `llaa33219/MicroT-test1-{1M..10K}`. Full analysis: [DATASET_COMPARISON_ANALYSIS.md](https://github.com/llaa33219/MicroMixer-4/blob/main/DATASET_COMPARISON_ANALYSIS.md).
327
+
328
+ ---
329
+
330
+ <div align="center">
331
+
332
+ [![GitHub](https://img.shields.io/badge/Back_to_Repository-MicroMixer--4-blue?style=for-the-badge&logo=github&color=%23007BFF)](https://github.com/llaa33219/MicroMixer-4)
333
+
334
+ <sub>Part of the <a href="https://github.com/llaa33219/MicroMixer-4">MicroMixer-4</a> research project — V87 Final (CCD-Mixer) family, 100K preset, UltraChat pretraining</sub>
335
+
336
+ </div>
fmsp_epoch_0.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f35e53bfe6a40454240a958bc3afa5b701e504595882e9a00609d677b4d3f319
3
+ size 413496
fmsp_epoch_1.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e33df950afa141d7c5af7df6bbe20d128033b17baad47c7291d70050b30c0747
3
+ size 413496
fmsp_epoch_2.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:23924dfb87478639e2664657b5dd4ea661c56f6496aebf00ccf1914d4e9cf694
3
+ size 413496
fmsp_epoch_3.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c40833b6c230a533ca663f295eb56c09ec05762c2158db2c3468fcf79d464bf1
3
+ size 413496
fmsp_epoch_4.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4370eaab979cb10003580de81a2ea36efd090de88a9534323ad3091f574deeab
3
+ size 413496
fmsp_epoch_5.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b7cb9c30c8071c2b850e51f9cc26b807aa270fef63556e745888faae113001cd
3
+ size 413496
fmsp_epoch_6.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bd522582126dd5282ffedd988050d348a01c4143fe9edd17ced273855d855330
3
+ size 413496
fmsp_epoch_7.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1cf2b94e53691deae765a5c2239cd6fa7b3086e5fa784a326661638d4308c524
3
+ size 413496
fmsp_epoch_8.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f80529448b10c27f305e6e4223d3b70e5e6af2b56df1613fa7ac568f2c233d03
3
+ size 413496
fmsp_epoch_9.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:01a4faab5fae31fa7426f7ed6170fd17b7e855b6514f87bfd147d36cf363664a
3
+ size 413496