grpo-qwen3-1p7b-math345 · end checkpoint (step 136)

Fine-tuned from Qwen/Qwen3-1.7B-Base with Vanilla GRPO against the dataset's ground-truth solutions.

  • Training data: q1716523669/MATH-Level345 (8,740 problems, MATH levels 3-5)
  • Eval: MATH-500 (500 prompts, pass@1, temperature 0.6)
  • This checkpoint: step 136 of 136 (2 epochs) — the end one
  • Its score: 0.6706

best is the step with the highest eval; end is step 136. Both are published because they answer different questions and the gap between them is a result in itself. The sibling repo is grpo-qwen3-1p7b-math345-best.

Hyperparameters

Effective batch 128 prompts per optimizer step, lr 3e-6, 2 epochs, 12 generations per prompt, 3072-token completions, temperature 1.0 for rollouts and 0.6 for eval, beta 0, bnpo loss, scale_rewards group, seed 42.

Eval curve

step MATH-500 pass@1
0 0.5437
10 0.6131
20 0.6290
30 0.6508
40 0.6706
50 0.6746
60 0.6766
70 0.6488
80 0.6845
90 0.6845
100 0.6607
110 0.6647
120 0.6667
130 0.6706

Files for audit

  • train.log — the complete training log this checkpoint came from
  • eval_curve.csv — the table above, machine-readable
  • run_config.json — resolved config as the trainer saw it

Code: williamium3000/trl-projects, branch n3-interaction-modes, projects/co-grpo-dp/.

Downloads last month
33
Safetensors
Model size
2B params
Tensor type
BF16
·
Video Preview
loading

Model tree for logan7000/grpo-qwen3-1p7b-math345-end

Finetuned
(408)
this model

Dataset used to train logan7000/grpo-qwen3-1p7b-math345-end