Add MicroT-test1-10K-UltraChat: FMSP model card + 10 epoch safetensors
Browse files- README.md +327 -0
- fmsp_epoch_0.safetensors +3 -0
- fmsp_epoch_1.safetensors +3 -0
- fmsp_epoch_2.safetensors +3 -0
- fmsp_epoch_3.safetensors +3 -0
- fmsp_epoch_4.safetensors +3 -0
- fmsp_epoch_5.safetensors +3 -0
- fmsp_epoch_6.safetensors +3 -0
- fmsp_epoch_7.safetensors +3 -0
- fmsp_epoch_8.safetensors +3 -0
- fmsp_epoch_9.safetensors +3 -0
README.md
ADDED
|
@@ -0,0 +1,327 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
library_name: pytorch
|
| 4 |
+
language:
|
| 5 |
+
- en
|
| 6 |
+
pipeline_tag: text-generation
|
| 7 |
+
tags:
|
| 8 |
+
- transformer
|
| 9 |
+
- byte-level
|
| 10 |
+
- causal-lm
|
| 11 |
+
- attention
|
| 12 |
+
- rope
|
| 13 |
+
- decoder-only
|
| 14 |
+
- fmsp
|
| 15 |
+
- micro-language-model
|
| 16 |
+
- sub-1m-parameters
|
| 17 |
+
model_name: MicroT-test1-10K-UltraChat
|
| 18 |
+
datasets:
|
| 19 |
+
- HuggingFaceH4/ultrachat_200k
|
| 20 |
+
- llaa33219/small-qa-en-10k
|
| 21 |
+
metrics:
|
| 22 |
+
- perplexity
|
| 23 |
+
- exact-match
|
| 24 |
+
---
|
| 25 |
+
|
| 26 |
+
<div align="center">
|
| 27 |
+
|
| 28 |
+
<img src="https://raw.githubusercontent.com/llaa33219/MicroMixer-4/main/logo.svg" width="300" alt="MicroMixer-4 Logo"/>
|
| 29 |
+
|
| 30 |
+
# MicroT-test1-10K-UltraChat
|
| 31 |
+
|
| 32 |
+
<img src="https://img.shields.io/badge/Parameters-9%2C808-blue?style=for-the-badge&logo=python&logoColor=white&color=%23007BFF" alt="Parameters"/>
|
| 33 |
+
<img src="https://img.shields.io/badge/Architecture-Transformer-orange?style=for-the-badge&color=%23FF6600" alt="Architecture"/>
|
| 34 |
+
<img src="https://img.shields.io/badge/Fine--tuning-FMSP-green?style=for-the-badge&color=%2300D620" alt="FMSP"/>
|
| 35 |
+
|
| 36 |
+
<br/>
|
| 37 |
+
<br/>
|
| 38 |
+
|
| 39 |
+
<table>
|
| 40 |
+
<tr>
|
| 41 |
+
<td align="center" style="padding: 20px;">
|
| 42 |
+
<strong style="color: #FF6600; font-size: 1.2em;">Micro Transformer — Reference Baseline</strong><br/><em>Vanilla Attention • RoPE • Byte-Level • Decoder-Only</em>
|
| 43 |
+
</td>
|
| 44 |
+
</tr>
|
| 45 |
+
</table>
|
| 46 |
+
|
| 47 |
+
[](https://github.com/llaa33219/MicroMixer-4)
|
| 48 |
+
|
| 49 |
+
</div>
|
| 50 |
+
|
| 51 |
+
<div style="background: linear-gradient(135deg, #FF660022, #007BFF22); padding: 20px; border-radius: 10px; border-left: 4px solid #FF6600;">
|
| 52 |
+
|
| 53 |
+
## 📋 Overview
|
| 54 |
+
|
| 55 |
+
**MicroT-test1-10K-UltraChat** is a **9,808-parameter** vanilla decoder-only **transformer** — multi-head causal self-attention with RoPE — pretrained on **UltraChat 200k** conversation data (instead of the project's Discord-Dialogues baseline) and then fine-tuned with **FMSP** on 9,012 general-knowledge QA pairs.
|
| 56 |
+
|
| 57 |
+
This is the **10K** member of the **MicroT-test1** family: the registered attention-based **reference baseline** of the MicroMixer-4 project, here rerun on UltraChat 200k as part of the **dataset-efficiency comparison study** — six parameter budgets × two architectures × two open pretraining corpora, all fine-tuned with the identical P05 FMSP recipe at seed 42. [Analysis](https://github.com/llaa33219/MicroMixer-4/blob/main/DATASET_COMPARISON_ANALYSIS.md).
|
| 58 |
+
|
| 59 |
+
It is deliberately boring — the standard 2018–2020 transformer recipe parameterized down to the sub-1M regime: no flash attention, no SwiGLU, no ALiBi, no QKNorm, no sliding window, no MQA/GQA, no MoE.
|
| 60 |
+
|
| 61 |
+
</div>
|
| 62 |
+
|
| 63 |
+
## 🏗️ Architecture
|
| 64 |
+
|
| 65 |
+
<div align="center">
|
| 66 |
+
|
| 67 |
+
```mermaid
|
| 68 |
+
graph TD
|
| 69 |
+
A[Byte Input] --> B[Embed 256→16]
|
| 70 |
+
B --> C[Transformer Block × 2]
|
| 71 |
+
C --> D[RMSNorm]
|
| 72 |
+
D --> E[LM Head Tied with Embed]
|
| 73 |
+
E --> F[Byte Output]
|
| 74 |
+
|
| 75 |
+
subgraph "Transformer Block (pre-norm)"
|
| 76 |
+
X[Input 16] --> N1[RMSNorm]
|
| 77 |
+
N1 --> AT["MHA 1 heads × d_head 16<br/>RoPE θ=10000 on q,k · causal SDPA"]
|
| 78 |
+
AT --> R1[+ residual]
|
| 79 |
+
R1 --> N2[RMSNorm]
|
| 80 |
+
N2 --> MLP["GELU MLP 16→56→16"]
|
| 81 |
+
MLP --> R2[+ residual]
|
| 82 |
+
end
|
| 83 |
+
|
| 84 |
+
style A fill:#007BFF,color:#fff
|
| 85 |
+
style F fill:#00D620,color:#fff
|
| 86 |
+
style AT fill:#FF6600,color:#fff
|
| 87 |
+
```
|
| 88 |
+
|
| 89 |
+
</div>
|
| 90 |
+
|
| 91 |
+
### Model Configuration
|
| 92 |
+
|
| 93 |
+
<table>
|
| 94 |
+
<tr>
|
| 95 |
+
<th style="background-color: #FF6600; color: white;">Parameter</th>
|
| 96 |
+
<th style="background-color: #007BFF; color: white;">Value</th>
|
| 97 |
+
</tr>
|
| 98 |
+
<tr><td>Hidden Dimension (d_model)</td><td><code>16</code></td></tr>
|
| 99 |
+
<tr><td>Attention Heads</td><td><code>1</code> (d_head = <code>16</code> at every size)</td></tr>
|
| 100 |
+
<tr><td>Number of Blocks</td><td><code>2</code></td></tr>
|
| 101 |
+
<tr><td>FFN Hidden</td><td><code>56</code></td></tr>
|
| 102 |
+
<tr><td>Position Encoding</td><td>RoPE θ=10000 on q/k only (non-persistent buffers)</td></tr>
|
| 103 |
+
<tr><td>Attention</td><td>Causal MHA via <code>F.scaled_dot_product_attention(is_causal=True)</code></td></tr>
|
| 104 |
+
<tr><td>Activation</td><td><code>GELU</code></td></tr>
|
| 105 |
+
<tr><td>Biases</td><td><b>None</b> — no bias parameters anywhere</td></tr>
|
| 106 |
+
<tr><td>Normalization</td><td><code>RMSNorm</code> (pre-norm)</td></tr>
|
| 107 |
+
<tr><td>Max Sequence Length</td><td><code>1024</code></td></tr>
|
| 108 |
+
<tr><td>Vocabulary Size</td><td><code>256</code> (byte-level)</td></tr>
|
| 109 |
+
<tr><td>Output Head</td><td>Tied with input embedding</td></tr>
|
| 110 |
+
</table>
|
| 111 |
+
|
| 112 |
+
### Core Components
|
| 113 |
+
|
| 114 |
+
```
|
| 115 |
+
┌──────────────────────────────────────────────┐
|
| 116 |
+
│ Transformer Block (×2) │
|
| 117 |
+
│ h = h + MHA(RMSNorm(h)) # RoPE q/k, causal│
|
| 118 |
+
│ h = h + MLP(RMSNorm(h)) # GELU d→ffn→d │
|
| 119 |
+
│ no biases, no flash, no tricks — vanilla │
|
| 120 |
+
└──────────────────────────────────────────────┘
|
| 121 |
+
```
|
| 122 |
+
|
| 123 |
+
The d_head=16 contract is hard-asserted across all six sizes so that attention-head behavior is comparable at every budget and never confounds the memorization measurements.
|
| 124 |
+
|
| 125 |
+
---
|
| 126 |
+
|
| 127 |
+
## 🎯 Generation Examples
|
| 128 |
+
|
| 129 |
+
<div style="background-color: #FF050515; padding: 15px; border-radius: 8px; border-left: 4px solid #FF6600;">
|
| 130 |
+
|
| 131 |
+
**Questions the model was trained on** (FMSP train set, 9,012 QA pairs — greedy decoding, `repetition_penalty=1.2`, `no_repeat_ngram_size=4`):
|
| 132 |
+
|
| 133 |
+
```
|
| 134 |
+
[Prompt] User: Who painted the Mona Lisa?
|
| 135 |
+
Assistant:
|
| 136 |
+
[Output] The stand is is the and the the the the the the the the the the the the the the the the the the…
|
| 137 |
+
```
|
| 138 |
+
|
| 139 |
+
<sub>degenerate — collapses into a repeating token loop; no real answer</sub>
|
| 140 |
+
|
| 141 |
+
```
|
| 142 |
+
[Prompt] User: Who painted The Starry Night?
|
| 143 |
+
Assistant:
|
| 144 |
+
[Output] The stand is is the and the and the the the the the the the the the the the the the the the the…
|
| 145 |
+
```
|
| 146 |
+
|
| 147 |
+
<sub>degenerate — collapses into a repeating token loop; no real answer</sub>
|
| 148 |
+
|
| 149 |
+
|
| 150 |
+
**Questions the model has never seen and cannot answer** (unanswerable probe — the correct behavior is to decline; the model's actual behavior is shown):
|
| 151 |
+
|
| 152 |
+
```
|
| 153 |
+
[Prompt] User: Who painted the Glimmering Frostberry?
|
| 154 |
+
Assistant:
|
| 155 |
+
[Output] The stand a th the the the the the the the the the the the the the the the the the the the the …
|
| 156 |
+
```
|
| 157 |
+
|
| 158 |
+
<sub>degenerate — collapses into a repeating token loop on out-of-distribution input</sub>
|
| 159 |
+
|
| 160 |
+
```
|
| 161 |
+
[Prompt] User: Who composed the Symphony of Hollow Dawn?
|
| 162 |
+
Assistant:
|
| 163 |
+
[Output] The stand is is the and the the the the the the the the the the the the the the the the the the…
|
| 164 |
+
```
|
| 165 |
+
|
| 166 |
+
<sub>degenerate — collapses into a repeating token loop on out-of-distribution input</sub>
|
| 167 |
+
|
| 168 |
+
|
| 169 |
+
</div>
|
| 170 |
+
|
| 171 |
+
---
|
| 172 |
+
|
| 173 |
+
## 📊 Results
|
| 174 |
+
|
| 175 |
+
<div style="background-color: #FF660015; padding: 15px; border-radius: 8px; border-left: 4px solid #FF6600;">
|
| 176 |
+
|
| 177 |
+
### Pretraining (UltraChat 200k, V76 recipe, 3 epochs)
|
| 178 |
+
|
| 179 |
+
| Metric | 1 ep | 2 ep | 3 ep |
|
| 180 |
+
|--------|------|------|------|
|
| 181 |
+
| Val PPL | 7.05 | 6.80 | **6.53** |
|
| 182 |
+
|
| 183 |
+
Identical recipe to the MicroMixer-4 mixer: AdamW lr 3e-3 · WSD (warmup 500) · wd 0.01 · bs 16 · seq 1024 · seed 42.
|
| 184 |
+
|
| 185 |
+
### FMSP fine-tuning (small-qa-en-10k, P05 recipe, 10 epochs)
|
| 186 |
+
|
| 187 |
+
| Metric | Value |
|
| 188 |
+
|--------|-------|
|
| 189 |
+
| Train QA pairs | 9,012 |
|
| 190 |
+
| Held-out QA pairs | 988 |
|
| 191 |
+
| Best-val checkpoint | `fmsp_epoch_9.safetensors` (val loss **1.7229**) |
|
| 192 |
+
| freeze_fraction | 0.05 (true freeze) |
|
| 193 |
+
| Loss | answer-only CE + probe KL (weight 0.5) |
|
| 194 |
+
|
| 195 |
+
### Evaluation battery (post-FMSP)
|
| 196 |
+
|
| 197 |
+
| Axis | MicroT-test1-10K-UltraChat |
|
| 198 |
+
|------|--------|
|
| 199 |
+
| Chatter fluency d2 (cycles) | 0.331 (9/9) |
|
| 200 |
+
| Full-988 EM (seed 42) | 0 |
|
| 201 |
+
| Q-relevance echo / hijack % | 0.0 / 0.0 |
|
| 202 |
+
| OOD hijack % | 0.0% |
|
| 203 |
+
| Unanswerable fabrication /18 | 0 |
|
| 204 |
+
| Discord PPL | 9.75 |
|
| 205 |
+
|
| 206 |
+
<sub>Single-seed run (seed 42); the discord-pretrained cards report a 3-seed mean for Full-988 EM. ‡ where marked: degenerate-pass.</sub>
|
| 207 |
+
|
| 208 |
+
### MicroT-test1 UltraChat family (same protocol, all sizes)
|
| 209 |
+
|
| 210 |
+
| Size | Params | 3ep Val PPL | Chatter d2 | Full-988 EM | qrel echo/hijack | OOD hijack |
|
| 211 |
+
|------|--------|-------------|------------|-------------|------------------|------------|
|
| 212 |
+
| 1M | 996,736 | 2.40 | 0.698 | 749 | 100.0 / 0.0 | 39.0% |
|
| 213 |
+
| 500K | 498,528 | 2.65 | 0.791 | 550 | 84.0 / 14.0 | 59.3% |
|
| 214 |
+
| 300K | 297,680 | 2.84 | 0.633 | 211 | 62.0 / 24.0 | 42.4% |
|
| 215 |
+
| 100K | 97,872 | 3.49 | 0.751 | 0 | 12.0 / 70.0 | 66.1% |
|
| 216 |
+
| 50K | 49,888 | 3.98 | 0.528 | 0 | 4.0 / 13.0 | 1.7% |
|
| 217 |
+
| **10K** | 9,808 | 6.53 | 0.331 | 0 | 0.0 / 0.0 | 0.0% |
|
| 218 |
+
|
| 219 |
+
<sub>Seed-42 single runs (pretrained on UltraChat 200k; the discord-pretrained families report 3-seed EM means).</sub>
|
| 220 |
+
|
| 221 |
+
</div>
|
| 222 |
+
|
| 223 |
+
---
|
| 224 |
+
|
| 225 |
+
## 📚 Training Data
|
| 226 |
+
|
| 227 |
+
<div style="background-color: #00D62015; padding: 15px; border-radius: 8px; border-left: 4px solid #00D620;">
|
| 228 |
+
|
| 229 |
+
1. **Pretraining**: [UltraChat 200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k) — 146K multi-turn conversations (`train_sft` split), flattened to `User:/Assistant:` format, 1024-byte sequences, 3 epochs.
|
| 230 |
+
2. **FMSP fine-tuning**: [small-qa-en-10k](https://huggingface.co/datasets/llaa33219/small-qa-en-10k) — 10K general-knowledge QA pairs (arts, science, history, geography, music…), split 9,012 train / 988 held-out. 10 epochs under the P05 recipe (5% of parameters frozen-true, answer-only CE, probe-KL 0.5).
|
| 231 |
+
|
| 232 |
+
</div>
|
| 233 |
+
|
| 234 |
+
---
|
| 235 |
+
|
| 236 |
+
## 🔧 Usage
|
| 237 |
+
|
| 238 |
+
### Files in this repository
|
| 239 |
+
- `fmsp_epoch_{0..9}.safetensors` — per-epoch FMSP weights (pickle-free safetensors). **`fmsp_epoch_9.safetensors` is the best-val checkpoint** for this size.
|
| 240 |
+
|
| 241 |
+
### Load and generate (local clone)
|
| 242 |
+
|
| 243 |
+
```python
|
| 244 |
+
import torch
|
| 245 |
+
from safetensors.torch import load_file
|
| 246 |
+
from src.model_v88_transformer import MicroMixerV88Transformer, v88_transformer_10k
|
| 247 |
+
from src.fmsp import attach_adapter
|
| 248 |
+
from src.tokenizer import ByteTokenizer
|
| 249 |
+
|
| 250 |
+
# Clone the code repository first:
|
| 251 |
+
# git clone https://github.com/llaa33219/MicroMixer-4.git && cd MicroMixer-4
|
| 252 |
+
|
| 253 |
+
cfg = v88_transformer_10k()
|
| 254 |
+
model = MicroMixerV88Transformer(cfg)
|
| 255 |
+
attach_adapter(model, d_model=cfg.d_model, rank=16) # FMSP adapter (trained weights are in the file)
|
| 256 |
+
model.load_state_dict(load_file("fmsp_epoch_9.safetensors"), strict=True)
|
| 257 |
+
model.eval()
|
| 258 |
+
|
| 259 |
+
tok = ByteTokenizer()
|
| 260 |
+
prompt = "User: Who painted the Mona Lisa?\n\nAssistant: "
|
| 261 |
+
ids = tok.encode(prompt)
|
| 262 |
+
if ids and ids[-1] == tok.eos_token_id:
|
| 263 |
+
ids = ids[:-1] # ByteTokenizer appends EOS; the prompt must end open
|
| 264 |
+
ids = torch.tensor([ids])
|
| 265 |
+
with torch.no_grad():
|
| 266 |
+
out = model.generate(
|
| 267 |
+
ids, max_new_tokens=200,
|
| 268 |
+
temperature=0.0, # greedy — used for all reported numbers
|
| 269 |
+
repetition_penalty=1.2,
|
| 270 |
+
no_repeat_ngram_size=4,
|
| 271 |
+
eos_token_id=tok.eos_token_id,
|
| 272 |
+
)
|
| 273 |
+
print(tok.decode(out[0].tolist()))
|
| 274 |
+
```
|
| 275 |
+
|
| 276 |
+
### Load from Hugging Face Hub (no clone of the weights needed)
|
| 277 |
+
|
| 278 |
+
```python
|
| 279 |
+
import torch
|
| 280 |
+
from huggingface_hub import hf_hub_download
|
| 281 |
+
from safetensors.torch import load_file
|
| 282 |
+
from src.model_v88_transformer import MicroMixerV88Transformer, v88_transformer_10k
|
| 283 |
+
from src.fmsp import attach_adapter
|
| 284 |
+
|
| 285 |
+
REPO = "llaa33219/MicroT-test1-10K-UltraChat"
|
| 286 |
+
|
| 287 |
+
cfg = v88_transformer_10k()
|
| 288 |
+
model = MicroMixerV88Transformer(cfg)
|
| 289 |
+
attach_adapter(model, d_model=cfg.d_model, rank=16)
|
| 290 |
+
model.load_state_dict(
|
| 291 |
+
load_file(hf_hub_download(REPO, "fmsp_epoch_9.safetensors")), strict=True)
|
| 292 |
+
model.eval()
|
| 293 |
+
# ... generate as above
|
| 294 |
+
```
|
| 295 |
+
|
| 296 |
+
---
|
| 297 |
+
|
| 298 |
+
## ⚠️ Limitations
|
| 299 |
+
|
| 300 |
+
<div style="background-color: #FF050515; padding: 15px; border-radius: 8px; border-left: 4px solid #FF0505;">
|
| 301 |
+
|
| 302 |
+
| Limitation | Description |
|
| 303 |
+
|------------|-------------|
|
| 304 |
+
| **Micro parameters** | 9,808 parameters; capacity is the binding constraint on every axis |
|
| 305 |
+
| **Reference baseline, not a product** | Exists to score the Mixer against attention at matched budget |
|
| 306 |
+
| **Knows only what it memorized** | Knowledge is limited to the 9,012 trained QA pairs + the pretraining-corpus distribution |
|
| 307 |
+
| **Does not abstain** | Unknown questions are answered with fabrication or degeneration, not refusal — see the examples above |
|
| 308 |
+
| **Byte-level noise** | 256-vocab byte tokenizer; PPL not comparable to BPE baselines |
|
| 309 |
+
| **Research use only** | Architecture/scaling research artifact, not a production model |
|
| 310 |
+
|
| 311 |
+
</div>
|
| 312 |
+
|
| 313 |
+
---
|
| 314 |
+
|
| 315 |
+
## 🧬 Context
|
| 316 |
+
|
| 317 |
+
This is the **10K** UltraChat-pretrained arm of the **dataset-efficiency comparison study** (24 arms: {V87, V88} × {UltraChat, SmolTalk2} × six sizes) in the [MicroMixer-4](https://github.com/llaa33219/MicroMixer-4) project. Each arm swaps the Discord-Dialogues pretraining baseline for an open corpus (UltraChat 200k) and reruns the identical P05 FMSP recipe at seed 42. Sibling repos: `llaa33219/MicroMixer-4-{1M..10K}-{UltraChat,SmolTalk2}` and `llaa33219/MicroT-test1-{1M..10K}-{UltraChat,SmolTalk2}`; the discord-pretrained baselines are `llaa33219/MicroMixer-4-{1M..10K}` and `llaa33219/MicroT-test1-{1M..10K}`. Full analysis: [DATASET_COMPARISON_ANALYSIS.md](https://github.com/llaa33219/MicroMixer-4/blob/main/DATASET_COMPARISON_ANALYSIS.md).
|
| 318 |
+
|
| 319 |
+
---
|
| 320 |
+
|
| 321 |
+
<div align="center">
|
| 322 |
+
|
| 323 |
+
[](https://github.com/llaa33219/MicroMixer-4)
|
| 324 |
+
|
| 325 |
+
<sub>Part of the <a href="https://github.com/llaa33219/MicroMixer-4">MicroMixer-4</a> research project — V88 transformer reference (MicroT-test1), 10K preset, UltraChat pretraining</sub>
|
| 326 |
+
|
| 327 |
+
</div>
|
fmsp_epoch_0.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:d95acd964c51c9a4e5aef98a761c7efa0145191362ba3acb2a149b3dfacf6824
|
| 3 |
+
size 43008
|
fmsp_epoch_1.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:632d7d87e6b22305dd32bb22e0c1f835b2e5d9159f447f7ec98e0db3f894d2dc
|
| 3 |
+
size 43008
|
fmsp_epoch_2.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:c0cb8742b2598847a9993637d2937b24b7692365991e25bf87e4df0964a6bb93
|
| 3 |
+
size 43008
|
fmsp_epoch_3.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:d16ae84ad821a8dbe3b3ff1a34380995c5a6fc235ed40c57000d17b2bad7994d
|
| 3 |
+
size 43008
|
fmsp_epoch_4.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:39e2ebf4a0710d002faab651eaebde76dcbab82fd9e46dd952b5daac390e6f78
|
| 3 |
+
size 43008
|
fmsp_epoch_5.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4651c5563492ec81641003fac348e0c10c01028f0488e9ebba56d2e345e4158e
|
| 3 |
+
size 43008
|
fmsp_epoch_6.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:8eb4497e0890313649307c17c719a4f71b554986bab9f8ec44d889372ab9abd5
|
| 3 |
+
size 43008
|
fmsp_epoch_7.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:5dab7cc42cacf85b4210e7aeb4b9aacf9d4da244f40c32aa32efbc769bcc8c46
|
| 3 |
+
size 43008
|
fmsp_epoch_8.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7d9bd14248c15f552e85845a0c95c8b02c5f17edab7575246847bda01b95907d
|
| 3 |
+
size 43008
|
fmsp_epoch_9.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7e5cfaa43195e5373083be6caa451b284624edbd7dbfa208dad1d8ba83d4de37
|
| 3 |
+
size 43008
|