logan7000/MATH-Level345
Viewer • Updated • 9.39k • 88
Fine-tuned from Qwen/Qwen3-1.7B-Base with Vanilla GRPO against the dataset's ground-truth solutions.
q1716523669/MATH-Level345 (8,740 problems, MATH levels 3-5)best is the step with the highest eval; end is step 136. Both are published
because they answer different questions and the gap between them is a result in
itself. The sibling repo is grpo-qwen3-1p7b-math345-end.
Effective batch 128 prompts per optimizer step, lr 3e-6, 2 epochs, 12
generations per prompt, 3072-token completions, temperature 1.0 for rollouts and
0.6 for eval, beta 0, bnpo loss, scale_rewards group, seed 42.
| step | MATH-500 pass@1 |
|---|---|
| 0 | 0.5437 |
| 10 | 0.6131 |
| 20 | 0.6290 |
| 30 | 0.6508 |
| 40 | 0.6706 |
| 50 | 0.6746 |
| 60 | 0.6766 |
| 70 | 0.6488 |
| 80 | 0.6845 |
| 90 | 0.6845 |
| 100 | 0.6607 |
| 110 | 0.6647 |
| 120 | 0.6667 |
| 130 | 0.6706 |
train.log — the complete training log this checkpoint came fromeval_curve.csv — the table above, machine-readablerun_config.json — resolved config as the trainer saw itCode: williamium3000/trl-projects,
branch n3-interaction-modes, projects/co-grpo-dp/.
Base model
Qwen/Qwen3-1.7B-Base