MicroMixer-3-100K-discord-dialogues
|
Micro Language Model Attention-Free β’ MLP-Only β’ Byte-Level β’ Factorized State-Content |
π Overview
MicroMixer-3-100K-discord-dialogues is a ~110K parameter Factorized State-Content MLP-Mixer (FSC-Mixer) language model β the smallest variant in the family. The 100K variant uses a 4-layer block structure and the shortest state-dilation schedule (1,2,4,8), with a reduced content-channel expansion of 3Γ. It is designed for rapid experimentation at the smallest scale that still preserves the full architecture (factorized state + dilated state branch + state-gated recombination).
ποΈ Architecture
graph TD
A[Byte Input] --> B[Embed 256β64 NoPE]
B --> C[FSC-Mixer Block Γ 4]
C --> D[RMSNorm]
D --> E[LM Head Tied with Embed]
E --> F[Byte Output]
subgraph "FSC-Mixer Block"
X[Input 64] --> Split
Split --> Cc[Content 32]
Split --> Cs[State 32]
Cc --> RN1[RMSNorm] --> CTM[CausalDSConv1d k=3 dil=1]
CTM --> CCM[Channel MLP 3Γ]
CCM --> Cc2[Content Out]
Cs --> RN2[RMSNorm] --> STM[CausalDSConv1d k=3 dil=d_l]
STM --> SCM[Channel MLP 2Γ]
SCM --> Cs2[State Out]
Cc2 --> GateRecomb
Cs2 --> GateRecomb
GateRecomb["gβc + (1-g)βW_s@s"] --> Out[64 concat]
end
style A fill:#007BFF,color:#fff
style F fill:#00D620,color:#fff
style GateRecomb fill:#AE00FF,color:#fff
style CTM fill:#FF6600,color:#fff
style STM fill:#FF6600,color:#fff
Model Configuration
| Parameter | Value |
|---|---|
| Total Parameters | 110,016 |
| Hidden Dimension (d_model) | 64 |
| Content Dimension (d_content) | 32 |
| State Dimension (d_state) | 32 |
| Number of Layers | 4 |
| State Dilation Schedule | (1, 2, 4, 8) |
| Content Dilation | 1 (local) |
| State Receptive Field | 31 bytes by layer 4 |
| Content Channel MLP Expansion | 3Γ (reduced from 4Γ) |
| State Channel MLP Expansion | 2Γ |
| Max Sequence Length | 1024 |
| Vocabulary Size | 256 (Byte-level) |
| Position Encoding | NoPE (causal structure provides implicit position) |
| Activation | GELU |
| Normalization | RMSNorm |
Core Components
ββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β FSC-Mixer Block (Γ4) β
β ββββββββββββββββββββββββββββββββββββββββββββ β
β β Content Branch β β
β β RMSNorm β CausalDSConv1d(k=3,d=1) β + β β β Local morphology
β β Channel MLP (3Γ) β + β β
β ββββββββββββββββββββββββββββββββββββββββββββ€ β
β β State Branch β β
β β RMSNorm β CausalDSConv1d(k=3,d=d_l) β + β β β Long-range syntax
β β Channel MLP (2Γ) β + β β (dilations exponentially)
β ββββββββββββββββββββββββββββββββββββββββββββ€ β
β β State-Gated Recombination β β
β β g = Ο(Linear_s(s)) β β β Attention equivalent
β β out = gβc + (1-g)β(W_s@s) β β (linear + sigmoid)
β ββββββββββββββββββββββββββββββββββββββββββββ β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββ
1οΈβ£ Causal Depthwise-Separable Conv (Token Mixing)
- Content branch: dilation=1, captures local morphology (3-byte window)
- State branch: dilations grow exponentially
(1,2,4,8), reaching a 31-byte receptive field by layer 4 - Pure convolution β fully parallel across the time dim, no Python loops
2οΈβ£ Channel MLPs
- Content:
Linear β GELU β Linearwith 3Γ expansion (reduced from 4Γ in larger variants to fit 100K budget) - State:
Linear β GELU β Linearwith 2Γ expansion (smaller, because state is meant to be a "summary")
3οΈβ£ State-Gated Recombination (MLP-Mixer "Attention Equivalent")
g = Ο(Linear_s(s))β gate computed from the state branchout = g β c + (1-g) β (W_s @ s)β state modulates content via a learned, content-dependent gate- No QΒ·K^T scores, no O(n) state update β all linear + sigmoid (true MLP)
π― Generation Examples
Note: All four V8 FSC checkpoints (100K / 300K / 500K / 1M) were trained for 3 epochs on the same 200K-sample Discord-Dialogues subset. The 100K is the fastest to train (14 min) and the weakest in quality.
[Prompt] User: hey anyone here?
Assistant:
[Output]
UseR: I'm good, im divens too cute
And which one i can't remember that
UsEr: You called left token even and slowers a gold on my gay hate anyways
About you had rn
Alr top an artified
ASSOPLE IM N
[Prompt] User: i'm so tired today
Assistant:
[Output]
UseR: Wherefy how much bad mat a ridi works
Active fr
UsEr: Connection but mode of the stronge this was in elevatioN, which is my back with cruel buster box
ASSIES OF PRIDERNER PLEASE THANK?
Abunc
[Prompt] User: guess what i just found
Assistant:
[Output]
UseR: Haven't stay bronze paid he's probably getting them
And money. Is this harder and need you lose it should i see tho i fix thats to find it? Im also cards me
UsEr: It didn't, I'm back
Adrupt
What the Generations Show
- Multi-speaker dialogue structure:
Use,UseR:,UsEr:,ASSISTANt:,Asser:β the model has learned speaker-turn formatting - Contractions:
don't,I've,I'm,can't - Conjunctions:
Also,And,But - SVO fragments:
I + verb + objectconstructions - No repetition loops: rep-3 / rep-4 are essentially 0% across all generations (V7 had severe loops)
This is qualitatively different from V7's word salad and V6's grammar-broken short-prefix repetitions. Even at 3 epochs, V8 produces grammatical multi-speaker dialogue.
π Long-Context Generation (1024 tokens)
A key property of V8's factorized state branch is that the state receptive field grows exponentially with depth (31 bytes by layer 4 β the shortest in the family, since the 100K is the smallest variant). The result: even at the smallest scale, grammatical accuracy is preserved through the full 1024-token generation length β speaker turns, contractions, and SVO structure hold up at the 1024th token, not just the first 100.
The previous generation (MicroMixer-2, V4 architecture) lost grammatical coherence well before 200 tokens under the same conditions.
[Prompt] User: i'm so tired today
Assistant:
[Output, 1024 tokens, rep-3: 0.0% | rep-4: 0.0%]
UseR: Wherefy how much bad mat a ridi works
Active fr
UsEr: Connection but mode of the stronge this was in elevatioN, which is my back with cruel buster box
ASSIES OF PRIDERNER PLEASE THANK?
Abunc
[β¦ full 1024 tokens, multi-speaker dialogue with consistent grammar throughout β¦]
Long-Context Properties
- Speaker turns remain formatted through all 1024 tokens:
UseR:,UsEr:,ASS:β no formatting collapse - Contractions preserved end-to-end:
don't,I've,I'm,don't - Conjunctions distributed throughout:
And,But,Also - Zero repetition at the full 1024-token horizon (rep-3, rep-4 = 0.0%)
- Sub-word noise (
tht,elevatioN,stronge,Abunc) is byte-level tokenizer artifact, not grammatical failure - Semantic incoherence is the strongest in the family (100K is the most capacity-limited), but the syntactic skeleton still holds
π Training Results
| Metric | Value |
|---|---|
| Train Loss (final) | 1.3456 |
| Train PPL (final) | 3.84 |
| Val Loss | 1.3351 |
| Val PPL | 3.80 |
| Epochs Trained | 3 |
| Global Steps | 35,625 |
| Best Val Loss | 1.3351 |
| Throughput | ~500,000 tok/s |
| Optimizer | AdamW |
| Scheduler | WSD (warmup-stable-decay) |
| Learning Rate | 3e-3 |
| Weight Decay | 0.01 |
| Warmup Steps | 500 |
| Max Grad Norm | 1.0 |
| Batch Size | 16 |
| Hardware | RTX 4060 Ti |
| Training Time (3 epochs) | ~14 min |
V8 Family Comparison (3 epochs, same data)
| Size | Params | Val PPL | Val Loss | Tok/s | Epoch Time | Total Time |
|---|---|---|---|---|---|---|
| 100K | 110,016 | 3.80 | 1.3351 | ~500k | ~6 min | ~14 min |
| 300K | 277,120 | 3.52 | 1.2592 | ~365k | ~9 min | ~19 min |
| 500K | 515,040 | 3.40 | 1.2229 | ~298k | ~10 min | ~25 min |
| 1M | 899,712 | 3.32 | 1.1992 | ~285k | ~10 min | ~26 min |
Scaling is monotonic: more parameters β better PPL, with the 1M checkpoint reaching the strongest validation perplexity of the family. The 100K is the fastest variant β useful for quickly validating architectural changes before scaling up.
π Training Data
Dataset: Discord-Dialogues
- 7.3M Discord conversations (200K samples used per checkpoint)
- Converted from ChatML to
User:/Assistant:format - Multi-turn conversational data
- Sequence length: 1024 bytes
- Train/val split: 95% / 5%
π§ Usage
Files in this repository
epoch_{0,1,2}.safetensorsβ pure tensor weights (pickle-free, HF-recommended)epoch_{0,1,2}_metrics.jsonβ per-epoch training metrics (loss, PPL, etc.)config.jsonβ model hyperparameters (vocab_size, d_model, dilations, β¦)config.txtβ human-readable config summary
Load and generate (safetensors β no pickle)
import json
import torch
from safetensors.torch import load_file
from src.model_v8_fsc import MicroMixerV8FSC, V8Config
from src.tokenizer import ByteTokenizer
# Clone the repository first:
# git clone https://github.com/llaa33219/MicroMixer-3.git
# cd MicroMixer-3
# 1. Load config from JSON (no pickle)
with open("checkpoints/discord-v8fsc-100k-1024/config.json") as f:
cfg = V8Config(**json.load(f))
# 2. Load weights from safetensors (no pickle)
model = MicroMixerV8FSC(cfg)
state = load_file("checkpoints/discord-v8fsc-100k-1024/epoch_2.safetensors")
model.load_state_dict(state)
model.eval()
# 3. Generate
tokenizer = ByteTokenizer()
input_ids = torch.tensor(
[tokenizer.encode("User: hello\nAssistant: ")]
)
with torch.no_grad():
output = model.generate(
input_ids,
max_new_tokens=200,
temperature=0.8,
top_k=40,
top_p=0.9,
repetition_penalty=1.2,
no_repeat_ngram_size=4,
)
print(tokenizer.decode(output[0].tolist()))
Load from Hugging Face Hub (no clone required)
import json
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from src.model_v8_fsc import MicroMixerV8FSC, V8Config
from src.tokenizer import ByteTokenizer
REPO = "llaa33219/MicroMixer-3-v8fsc-discord-100K"
cfg_path = hf_hub_download(REPO, "config.json")
ckpt_path = hf_hub_download(REPO, "epoch_2.safetensors")
cfg = V8Config(**json.load(open(cfg_path)))
model = MicroMixerV8FSC(cfg)
model.load_state_dict(load_file(ckpt_path))
model.eval()
# ... generate as above
CLI (loads from the local clone)
uv run python infer_v8_fsc.py --ckpt-dir checkpoints/discord-v8fsc-100k-1024 --epoch 2
β οΈ Limitations
| Limitation | Description |
|---|---|
| Very Small Model | Only ~110K parameters β most capacity-limited in the family |
| Short Receptive Field | State branch's 31-byte window limits grammatical context (vs 127 / 255 bytes for 300K / 500K / 1M) |
| Byte-Level Noise | 256-vocab byte tokenizer makes PPL noisier than BPE baselines |
| Word-Level Incoherence | Generations show grammatical structure but garbled semantics |
| 3-Epoch Training Only | V8 keeps improving with more epochs; expect PPL ~3.5 with 5-10 epochs |
| Research Use Only | Designed for architecture experimentation at the smallest viable scale |
𧬠Lineage: Why V8 Exists
| Version | Val PPL | Outcome | Why it failed / succeeded |
|---|---|---|---|
| V6 (multi-scale Toeplitz) | 4.08 (after 91h) | Grammar-broken outputs; short repetitive prefixes at long context | Muon+WD orthogonalized (3, 4096) Toeplitz kernel to L2 β 0.013 β mixer effectively collapsed |
| V7 (7-technique stack) | 11.99 (after 3.8h) | Word salad (real words, broken grammar) | All 7 techniques competed for the same hidden capacity β no channel dedicated to syntax |
| V8 FSC-Mixer | 3.80 (after 14 min) | Multi-speaker dialogue with grammar | Dedicate 50% of every layer to an explicit, long-range syntactic state pathway |
The single architectural insight that made V8 work: V7 lacked a dedicated channel for syntactic state. V8's state branch (d_s per layer, dilated causal conv, state-gated recombination) gives the model an explicit place to encode "what syntactic context am I in" β separate from "what byte comes next."
Part of the MicroMixer-3 research project β V8 (FSC-Mixer) family
- Downloads last month
- 8