qwen3-4b-gdn-otmix-2p0b-opd

Model weights of the Qwen3-4B GDN-hybrid on-policy-distillation (OPD) WSD ladder run qwen-wsd-otmix-h32k-2p0b-v1 (W&B id qwen-otmix-h32k-b200-2p0b-v1-20260904) at step 11701 / 2,000,169,520 consumed tokens, H=32K.

Ladder: OpenThoughts prompt mix (stage3_prompts_v1, ~15% RUG / ~20% general), horizon 32,768, 4x B200 (3-GPU ZeRO-1 trainer + 1 vLLM sampler, gen_batch 256). Each rung = ~200M flat-LR tokens at 2e-5 plus an 839-step (110M-token) linear decay, seeded from the previous rung's pre-decay end-minus-0839 full state; the first rung seeds from wsd-flat-ext800-predecay-890M (pinkskin/qwen3-4b-gdn-wsd-ladder).

LR tail: decayed to 0. Compare rungs with matching tails.

Decayed final of the wsd_qwen_otmix_h32k_b200x4_2p0b rung; the pre-decay full state (end-minus-0839) is kept locally and not uploaded.

Custom GDN-hybrid architecture - register before loading (see project repo).

Source Git commit: d86fbef09d35f4e4d7943ec51d2b3732eb1fed46

Downloads last month
9
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for arianraje/qwen3-4b-gdn-otmix-2p0b-opd

Finetuned
Qwen/Qwen3-4B
Finetuned
(992)
this model