Qwen 3 4B RLM RLVR
Everything from my experiments on training RLMs.
4B • Updated • 3Note SFT attempt, per-root-turn trajectory format
lsteno/Qwen3-4B-Instruct-2507-RLM-RL-depth1-r4-a8-lr5e-7-s150-lora
Updated • 9Note LoRA RL, depth-1, rank 4 alpha 8, LR 5e-7, 150 steps
lsteno/Qwen3-4B-Instruct-2507-RLM-RL-depth1-r4-a8-lr1e-5-s150-lora
Updated • 3Note LoRA RL, depth-1, rank 4 alpha 8, LR 1e-5, 150 steps
lsteno/qwen3-rlm-depth1-r4-a8-lr1e-4-s150-bal35f40v1-lora
Updated • 4Note LoRA RL, balanced BEEG, rank 4 alpha 8, LR 1e-4, 150 steps
lsteno/qwen3-rlm-depth1-r16-a32-lr5e-7-s150-bal35f40v1-lora
Updated • 3Note LoRA RL, balanced BEEG, rank 16 alpha 32, LR 5e-7, 150 steps
lsteno/qwen3-rlm-depth1-r16-a32-lr1e-5-s150-bal35f40v1-lora
Updated • 2Note LoRA RL, balanced BEEG, rank 16 alpha 32, LR 1e-5, 150 steps
lsteno/qwen3-rlm-depth1-r16-a32-lr1e-4-s150-bal35f40v1-lora
Updated • 2Note LoRA RL, balanced BEEG, rank 16 alpha 32, LR 1e-4, 150 steps
lsteno/qwen3-rlm-depth1-r64-a128-lr5e-7-s150-bal35f40v1-lora
Updated • 8Note LoRA RL, balanced BEEG, rank 64 alpha 128, LR 5e-7, 150 steps
lsteno/qwen3-rlm-depth1-r64-a128-lr1e-5-s150-bal35f40v1-lora
Updated • 2Note LoRA RL, balanced BEEG, rank 64 alpha 128, LR 1e-5, 150 steps
lsteno/Qwen3-4B-Instruct-2507-RLM-RLVR-FullFT-lr5e-6-depth1-v1
Text Generation • 4B • Updated • 18Note Full fine-tune RLVR, depth-1, LR 5e-6, 150 steps
lsteno/Qwen3-4B-Instruct-2507-RLM-RLVR-FullFT-lr1e-5-depth1-v1
4B • Updated • 3Note Full fine-tune RLVR, depth-1, LR 1e-5, 150 steps
lsteno/BEEG-agents
Viewer • Updated • 3.02k • 48Note BEEG-agents dataset used for RLM RLVR train/eval splits
lsteno/Qwen3-4B-Instruct-2507-RLM-RLVR-depth2-recursive-r64-a128-lr1e-5-adapter
Reinforcement Learning • Updated • 3Note Depth-2 recursive RLM RLVR LoRA adapter, r64/a128/lr1e-5, step 150.