qwen3-4b-gdn-otmix-2p0b-opd
Model weights of the Qwen3-4B GDN-hybrid on-policy-distillation (OPD) WSD ladder run qwen-wsd-otmix-h32k-2p0b-v1
(W&B id qwen-otmix-h32k-b200-2p0b-v1-20260904) at step 11701 / 2,000,169,520 consumed tokens, H=32K.
Ladder: OpenThoughts prompt mix (stage3_prompts_v1, ~15% RUG / ~20% general), horizon 32,768, 4x B200 (3-GPU ZeRO-1 trainer + 1 vLLM sampler, gen_batch 256). Each rung = ~200M flat-LR tokens at 2e-5 plus an 839-step (110M-token) linear decay, seeded from the previous rung's pre-decay end-minus-0839 full state; the first rung seeds from wsd-flat-ext800-predecay-890M (pinkskin/qwen3-4b-gdn-wsd-ladder).
LR tail: decayed to 0. Compare rungs with matching tails.
Decayed final of the wsd_qwen_otmix_h32k_b200x4_2p0b rung; the pre-decay full state (end-minus-0839) is kept locally and not uploaded.
Custom GDN-hybrid architecture - register before loading (see project repo).
Source Git commit: d86fbef09d35f4e4d7943ec51d2b3732eb1fed46
- Downloads last month
- 9