qwen3-4b-gdn-mathmix-1p2b-opd

Model weights of the Qwen3-4B GDN-hybrid on-policy-distillation (OPD) math-mix WSD ladder run qwen-wsd-mathmix-h16k-1p2b-v1 (W&B id qwen-mathmix-h16k-empire-1p2b-v1-20260831) at step 7205 / 1,200,015,743 consumed tokens, H=16K.

LR tail: decayed to 1e-6. The mainline mathmix ladder decayed to 1e-6 through the 1.2B rung; from the 1.4B rung onward the recipe decays to 0 (user-authorized change, 2026-09-01). Compare rungs with matching tails.

Decayed final of the 1.2B rung. Battery (slurm 58126-58131, 2026-09-02, bundled here as JSONs): AIME24 43.3 / 73.3 and AIME25 40.8 / 63.3 pass@1 / pass@8 at 32K; MATH-500 think pass@1 91.4.

Custom GDN-hybrid architecture - register before loading (see project repo).

Source Git commit: d86fbef09d35f4e4d7943ec51d2b3732eb1fed46

Downloads last month
21
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for arianraje/qwen3-4b-gdn-hybrid-1.2B-OPD

Finetuned
Qwen/Qwen3-4B
Finetuned
(1010)
this model