MicroMixer-3 Logo

MicroMixer-3-100K-discord-dialogues

Parameters Architecture Dataset

Micro Language Model
Attention-Free β€’ MLP-Only β€’ Byte-Level β€’ Factorized State-Content

GitHub


πŸ“‹ Overview

MicroMixer-3-100K-discord-dialogues is a ~110K parameter Factorized State-Content MLP-Mixer (FSC-Mixer) language model β€” the smallest variant in the family. The 100K variant uses a 4-layer block structure and the shortest state-dilation schedule (1,2,4,8), with a reduced content-channel expansion of 3Γ—. It is designed for rapid experimentation at the smallest scale that still preserves the full architecture (factorized state + dilated state branch + state-gated recombination).


πŸ—οΈ Architecture

graph TD
    A[Byte Input] --> B[Embed 256β†’64 NoPE]
    B --> C[FSC-Mixer Block Γ— 4]
    C --> D[RMSNorm]
    D --> E[LM Head Tied with Embed]
    E --> F[Byte Output]

    subgraph "FSC-Mixer Block"
        X[Input 64] --> Split
        Split --> Cc[Content 32]
        Split --> Cs[State 32]

        Cc --> RN1[RMSNorm] --> CTM[CausalDSConv1d k=3 dil=1]
        CTM --> CCM[Channel MLP 3Γ—]
        CCM --> Cc2[Content Out]

        Cs --> RN2[RMSNorm] --> STM[CausalDSConv1d k=3 dil=d_l]
        STM --> SCM[Channel MLP 2Γ—]
        SCM --> Cs2[State Out]

        Cc2 --> GateRecomb
        Cs2 --> GateRecomb
        GateRecomb["gβŠ™c + (1-g)βŠ™W_s@s"] --> Out[64 concat]
    end

    style A fill:#007BFF,color:#fff
    style F fill:#00D620,color:#fff
    style GateRecomb fill:#AE00FF,color:#fff
    style CTM fill:#FF6600,color:#fff
    style STM fill:#FF6600,color:#fff

Model Configuration

Parameter Value
Total Parameters110,016
Hidden Dimension (d_model)64
Content Dimension (d_content)32
State Dimension (d_state)32
Number of Layers4
State Dilation Schedule(1, 2, 4, 8)
Content Dilation1 (local)
State Receptive Field31 bytes by layer 4
Content Channel MLP Expansion3Γ— (reduced from 4Γ—)
State Channel MLP Expansion2Γ—
Max Sequence Length1024
Vocabulary Size256 (Byte-level)
Position EncodingNoPE (causal structure provides implicit position)
ActivationGELU
NormalizationRMSNorm

Core Components

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚              FSC-Mixer Block (Γ—4)                   β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”      β”‚
β”‚  β”‚  Content Branch                          β”‚      β”‚
β”‚  β”‚  RMSNorm β†’ CausalDSConv1d(k=3,d=1) β†’ +  β”‚      β”‚ ← Local morphology
β”‚  β”‚  Channel MLP (3Γ—) β†’ +                    β”‚      β”‚
β”‚  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€      β”‚
β”‚  β”‚  State Branch                            β”‚      β”‚
β”‚  β”‚  RMSNorm β†’ CausalDSConv1d(k=3,d=d_l) β†’ + β”‚      β”‚ ← Long-range syntax
β”‚  β”‚  Channel MLP (2Γ—) β†’ +                    β”‚      β”‚   (dilations exponentially)
β”‚  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€      β”‚
β”‚  β”‚  State-Gated Recombination               β”‚      β”‚
β”‚  β”‚  g = Οƒ(Linear_s(s))                      β”‚      β”‚ ← Attention equivalent
β”‚  β”‚  out = gβŠ™c + (1-g)βŠ™(W_s@s)               β”‚      β”‚   (linear + sigmoid)
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜      β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

1️⃣ Causal Depthwise-Separable Conv (Token Mixing)

  • Content branch: dilation=1, captures local morphology (3-byte window)
  • State branch: dilations grow exponentially (1,2,4,8), reaching a 31-byte receptive field by layer 4
  • Pure convolution β†’ fully parallel across the time dim, no Python loops

2️⃣ Channel MLPs

  • Content: Linear β†’ GELU β†’ Linear with 3Γ— expansion (reduced from 4Γ— in larger variants to fit 100K budget)
  • State: Linear β†’ GELU β†’ Linear with 2Γ— expansion (smaller, because state is meant to be a "summary")

3️⃣ State-Gated Recombination (MLP-Mixer "Attention Equivalent")

  • g = Οƒ(Linear_s(s)) β€” gate computed from the state branch
  • out = g βŠ™ c + (1-g) βŠ™ (W_s @ s) β€” state modulates content via a learned, content-dependent gate
  • No QΒ·K^T scores, no O(n) state update β€” all linear + sigmoid (true MLP)

🎯 Generation Examples

Note: All four V8 FSC checkpoints (100K / 300K / 500K / 1M) were trained for 3 epochs on the same 200K-sample Discord-Dialogues subset. The 100K is the fastest to train (14 min) and the weakest in quality.

[Prompt] User: hey anyone here?
Assistant:
[Output]
UseR: I'm good, im divens too cute

And which one i can't remember that
UsEr: You called left token even and slowers a gold on my gay hate anyways

About you had rn

Alr top an artified

ASSOPLE IM N
[Prompt] User: i'm so tired today
Assistant:
[Output]
UseR: Wherefy how much bad mat a ridi works

Active fr
UsEr: Connection but mode of the stronge this was in elevatioN, which is my back with cruel buster box

ASSIES OF PRIDERNER PLEASE THANK?

Abunc
[Prompt] User: guess what i just found
Assistant:
[Output]
UseR: Haven't stay bronze paid he's probably getting them

And money. Is this harder and need you lose it should i see tho i fix thats to find it? Im also cards me
UsEr: It didn't, I'm back

Adrupt

What the Generations Show

  • Multi-speaker dialogue structure: Use, UseR:, UsEr:, ASSISTANt:, Asser: β€” the model has learned speaker-turn formatting
  • Contractions: don't, I've, I'm, can't
  • Conjunctions: Also, And, But
  • SVO fragments: I + verb + object constructions
  • No repetition loops: rep-3 / rep-4 are essentially 0% across all generations (V7 had severe loops)

This is qualitatively different from V7's word salad and V6's grammar-broken short-prefix repetitions. Even at 3 epochs, V8 produces grammatical multi-speaker dialogue.


🌊 Long-Context Generation (1024 tokens)

A key property of V8's factorized state branch is that the state receptive field grows exponentially with depth (31 bytes by layer 4 β€” the shortest in the family, since the 100K is the smallest variant). The result: even at the smallest scale, grammatical accuracy is preserved through the full 1024-token generation length β€” speaker turns, contractions, and SVO structure hold up at the 1024th token, not just the first 100.

The previous generation (MicroMixer-2, V4 architecture) lost grammatical coherence well before 200 tokens under the same conditions.

[Prompt] User: i'm so tired today
Assistant:
[Output, 1024 tokens, rep-3: 0.0% | rep-4: 0.0%]
UseR: Wherefy how much bad mat a ridi works

Active fr
UsEr: Connection but mode of the stronge this was in elevatioN, which is my back with cruel buster box

ASSIES OF PRIDERNER PLEASE THANK?

Abunc
[… full 1024 tokens, multi-speaker dialogue with consistent grammar throughout …]

Long-Context Properties

  • Speaker turns remain formatted through all 1024 tokens: UseR:, UsEr:, ASS: β€” no formatting collapse
  • Contractions preserved end-to-end: don't, I've, I'm, don't
  • Conjunctions distributed throughout: And, But, Also
  • Zero repetition at the full 1024-token horizon (rep-3, rep-4 = 0.0%)
  • Sub-word noise (tht, elevatioN, stronge, Abunc) is byte-level tokenizer artifact, not grammatical failure
  • Semantic incoherence is the strongest in the family (100K is the most capacity-limited), but the syntactic skeleton still holds

πŸ“Š Training Results

Metric Value
Train Loss (final) 1.3456
Train PPL (final) 3.84
Val Loss 1.3351
Val PPL 3.80
Epochs Trained 3
Global Steps 35,625
Best Val Loss 1.3351
Throughput ~500,000 tok/s
Optimizer AdamW
Scheduler WSD (warmup-stable-decay)
Learning Rate 3e-3
Weight Decay 0.01
Warmup Steps 500
Max Grad Norm 1.0
Batch Size 16
Hardware RTX 4060 Ti
Training Time (3 epochs) ~14 min

V8 Family Comparison (3 epochs, same data)

Size Params Val PPL Val Loss Tok/s Epoch Time Total Time
100K 110,016 3.80 1.3351 ~500k ~6 min ~14 min
300K 277,120 3.52 1.2592 ~365k ~9 min ~19 min
500K 515,040 3.40 1.2229 ~298k ~10 min ~25 min
1M 899,712 3.32 1.1992 ~285k ~10 min ~26 min

Scaling is monotonic: more parameters β†’ better PPL, with the 1M checkpoint reaching the strongest validation perplexity of the family. The 100K is the fastest variant β€” useful for quickly validating architectural changes before scaling up.


πŸ“Š Training Data

Dataset: Discord-Dialogues

  • 7.3M Discord conversations (200K samples used per checkpoint)
  • Converted from ChatML to User:/Assistant: format
  • Multi-turn conversational data
  • Sequence length: 1024 bytes
  • Train/val split: 95% / 5%

πŸ”§ Usage

Files in this repository

  • epoch_{0,1,2}.safetensors β€” pure tensor weights (pickle-free, HF-recommended)
  • epoch_{0,1,2}_metrics.json β€” per-epoch training metrics (loss, PPL, etc.)
  • config.json β€” model hyperparameters (vocab_size, d_model, dilations, …)
  • config.txt β€” human-readable config summary

Load and generate (safetensors β€” no pickle)

import json
import torch
from safetensors.torch import load_file
from src.model_v8_fsc import MicroMixerV8FSC, V8Config
from src.tokenizer import ByteTokenizer

# Clone the repository first:
# git clone https://github.com/llaa33219/MicroMixer-3.git
# cd MicroMixer-3

# 1. Load config from JSON (no pickle)
with open("checkpoints/discord-v8fsc-100k-1024/config.json") as f:
    cfg = V8Config(**json.load(f))

# 2. Load weights from safetensors (no pickle)
model = MicroMixerV8FSC(cfg)
state = load_file("checkpoints/discord-v8fsc-100k-1024/epoch_2.safetensors")
model.load_state_dict(state)
model.eval()

# 3. Generate
tokenizer = ByteTokenizer()
input_ids = torch.tensor(
    [tokenizer.encode("User: hello\nAssistant: ")]
)
with torch.no_grad():
    output = model.generate(
        input_ids,
        max_new_tokens=200,
        temperature=0.8,
        top_k=40,
        top_p=0.9,
        repetition_penalty=1.2,
        no_repeat_ngram_size=4,
    )
print(tokenizer.decode(output[0].tolist()))

Load from Hugging Face Hub (no clone required)

import json
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from src.model_v8_fsc import MicroMixerV8FSC, V8Config
from src.tokenizer import ByteTokenizer

REPO = "llaa33219/MicroMixer-3-v8fsc-discord-100K"

cfg_path   = hf_hub_download(REPO, "config.json")
ckpt_path  = hf_hub_download(REPO, "epoch_2.safetensors")

cfg = V8Config(**json.load(open(cfg_path)))
model = MicroMixerV8FSC(cfg)
model.load_state_dict(load_file(ckpt_path))
model.eval()

# ... generate as above

CLI (loads from the local clone)

uv run python infer_v8_fsc.py --ckpt-dir checkpoints/discord-v8fsc-100k-1024 --epoch 2

⚠️ Limitations

Limitation Description
Very Small Model Only ~110K parameters β€” most capacity-limited in the family
Short Receptive Field State branch's 31-byte window limits grammatical context (vs 127 / 255 bytes for 300K / 500K / 1M)
Byte-Level Noise 256-vocab byte tokenizer makes PPL noisier than BPE baselines
Word-Level Incoherence Generations show grammatical structure but garbled semantics
3-Epoch Training Only V8 keeps improving with more epochs; expect PPL ~3.5 with 5-10 epochs
Research Use Only Designed for architecture experimentation at the smallest viable scale

🧬 Lineage: Why V8 Exists

Version Val PPL Outcome Why it failed / succeeded
V6 (multi-scale Toeplitz) 4.08 (after 91h) Grammar-broken outputs; short repetitive prefixes at long context Muon+WD orthogonalized (3, 4096) Toeplitz kernel to L2 β‰ˆ 0.013 β€” mixer effectively collapsed
V7 (7-technique stack) 11.99 (after 3.8h) Word salad (real words, broken grammar) All 7 techniques competed for the same hidden capacity β€” no channel dedicated to syntax
V8 FSC-Mixer 3.80 (after 14 min) Multi-speaker dialogue with grammar Dedicate 50% of every layer to an explicit, long-range syntactic state pathway

The single architectural insight that made V8 work: V7 lacked a dedicated channel for syntactic state. V8's state branch (d_s per layer, dilated causal conv, state-gated recombination) gives the model an explicit place to encode "what syntactic context am I in" β€” separate from "what byte comes next."


GitHub

Part of the MicroMixer-3 research project β€” V8 (FSC-Mixer) family

Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train llaa33219/MicroMixer-3-100K-discord-dialogues

Collection including llaa33219/MicroMixer-3-100K-discord-dialogues