MiMo-7B GDN Hybrid โ€” stage3-1.8B-OPD

Decayed final of the 1.8B-token rung of the MiMo on-policy-distillation (OPD) WSD ladder (1:4 uniform hybrid of Qwen2 attention and Gated DeltaNet layers, teacher XiaomiMiMo/MiMo-7B-RL-0530). Run mimo-mathmix-h32k-runpod-1p8b-v1-20260903; snapshots/final, step 11632, 1,800,027,023 consumed generation tokens; model window 65,536.

Recipe: extends arianraje/mimo-7b-gdn-hybrid-1.6B-OPD from its pre-decay state (1,521.4M tokens) with TWO changes vs that row. Deliberate: rollout horizon horizon_max 16,384 โ†’ 32,768 (reasoning rollouts trained at H=32K for the whole rung). Forced by trainer memory: sampler max_model_len 65,536 โ†’ 49,152, whose only effect is that 32K-prompt retrieval (RUG) items have their generation clamped to ~16K โ€” the budget they had on every H=16K rung. Everything else (stage-3 math mixture, decay 600 steps, queue 6.0) is inherited. The pre-decay trunk of this rung is public at arianraje/mimo-7b-gdn-opd-predecay-1721m-step11144.

Evaluation (TABLES.md protocol, raw JSONs under full_eval/)

Metric 1.6B-OPD 1.8B-OPD
AIME24 think pass@1 / pass@8 @32K (n=30ร—8) 62.1 / 83.3 63.3 / 76.7
AIME25 think pass@1 / pass@8 @32K (n=30ร—8) 47.5 / 73.3 49.2 / 73.3
MATH-500 think pass@1 @32K 93.8 93.4
MATH-500 no-think pass@1 @4K 68.2 70.2
GSM8K no-think strict / flexible @1K 57.6 / 63.8 58.5 / 65.2
MMLU 5-shot 54.2 53.8
PIQA / HellaSwag / ARC-E / ARC-C / Winogrande 72.6 / 60.3 / 63.6 / 39.1 / 59.3 72.0 / 60.3 / 62.8 / 39.6 / 59.7
NIAH multikey 4K / 8K / 16K / 32K (n=500) 98.6 / 98.0 / 96.6 / 89.4 96.8 / 97.2 / 95.0 / 84.6
NIAH single, multiquery (all lengths) โ‰ฅ99.8 โ‰ฅ99.8
trunc / mean gen tokens: MATH-500, AIME24, AIME25 3.4%/8169, 30.4%/22273, 35.8%/23806 3.6%/8168, 26.7%/21925, 32.5%/23552

AIME cells are n=30 problems (ยฑ~14 pts on pass@1) โ€” trend data; MATH-500 (n=500) is the powered cell. full_eval/*.generations.json carry every sample.

Loading

Custom architecture; needs the mimo_gdn model registration from the training repo (transformers 4.57.x + flash-linear-attention 0.5.x), same as the other arianraje/mimo-7b-gdn-hybrid-*-OPD checkpoints.

Downloads last month
524
Safetensors
Model size
8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for arianraje/mimo-7b-gdn-hybrid-1.8B-OPD

Finetuned
(15)
this model