mimo-7b-gdn-hybrid-1B-OPD

Full state (BF16 model shards + FP32 optimizer shards + trainer_state) of the MiMo-7B GDN-hybrid OPD run mimo-wsd-rugfix-h200x3-ext800-decay0-v1 (W&B id mimo-rugfix-h200x3-ext800-decay0-v1-20260816) at step 6454 / 1,000,002,516 consumed tokens, H=32K. Saved on the Empire 3xH200 7-day reservation (job 16007, WSD OPD scaling ladder). This is the decayed ext800 (1B) final state at the rung endpoint.

Custom mimo_gdn architecture - register before loading (see project repo).

Source Git commit: e3cc4c0ff85a10b0660af706e80f21bfb252677b

Downloads last month
2,992
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for arianraje/mimo-7b-gdn-hybrid-1B-OPD

Finetuned
(15)
this model