Gemma-3-12B - GT-Reward on MMR1 (mmupt recipe)

GRPO with ground-truth rewards, MMR1-Math-RL-Data-v0, 481 steps (mmupt recipe: beta=0.01, 10 generations/prompt, cap 2048, T=0.7, 12 prompts/update).

folder step note
best/ 100 best in-loop val (eval_reward = 0.375)
endpoint/ 481 end of training

Training note: five gradient spikes occurred (max grad-norm ~1.0e6, clipped at 1.0); best-by-val fell early (s100) and the in-loop curve was flat afterwards. Step-count fingerprint: 481 = mmupt recipe (the older recipe runs 722 steps). training/ holds metrics, train.log, and the Slurm joblog.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for logan7000/mllm-mmr1-gt-gemma3-12b-mmupt-full

Finetuned
(397)
this model