Teaching Gemma 3 chain-of-thought reasoning
The short version
For the Google Tunix Hackathon I fine-tuned Gemma 3 1B to reason step by step before answering. The recipe is two supervised fine-tuning passes and a reinforcement stage with Group Relative Policy Optimization, all built on Tunix, Google's JAX post-training library, and trained on TPU v5e. The point was not a leaderboard number. It was to build each stage by hand so I understand where reasoning actually comes from. This is the reinforcement-learning member of my series, next to the from-scratch gpt2-nano and the applied Phi-4 work.
The Tunix hackathon that started it
The Google Tunix Hackathon asked people to build with Tunix, a JAX library for post-training. That was a good excuse to stop reading about reasoning models and train one. I picked mathematics as the domain, because a final numeric answer is either correct or not. That gives an exact reward with no room for a model to talk its way to a good score, which is exactly what the reinforcement stage needs.
Reasoning is trained, not prompted
A model does not become a reasoner because you ask it nicely in the prompt. The public recipe has three moving parts. First you supervise the format, so the model learns to show its work. Then you supervise the domain, so the work is actually correct. Then you run a reinforcement stage where the reward is checked mechanically rather than learned from a preference model. Each part does a different job, and you only see that clearly once you run them in order.
Teaching the format, then the math
The first supervised pass teaches the chat and reasoning template. After it, the model emits an explicit reasoning segment followed by a delimited final answer. It has the shape of reasoning even when the content is thin.
The second supervised pass fine-tunes on worked solutions. Now the reasoning segment carries real mathematical steps instead of just the right structure. The order matters. Format first gives the later stages a stable thing to reward.
GRPO and a reward functions
The third stage is GRPO. It samples a group of completions per prompt, scores each one with the reward function, and shifts the policy toward the completions that beat the group mean. The group itself is the baseline, which is what lets GRPO drop the learned value network that other methods carry. For a group of sampled completions with rewards , each advantage is the group-normalized reward
The reward is a sum of verifiable parts. Exact-match on the final answer, a format term that checks the reasoning and answer delimiters are present and well formed, and a soft term for producing a parseable number at all. Every part is computed by a checker rather than a model, so fluent but wrong text earns nothing.
What the stages actually changed
The clear effect of the GRPO stage is that completions converge on the required reasoning-then-answer structure and stop emitting bare answers, which is the format term of the reward doing its job. At the same time the exact-match term pulls the content toward correct final answers on the training distribution. The contribution here is the working pipeline that produces that behavior, not a state-of-the-art score.
What building the recipe taught me
The reward is the design. Once the reward is a set of checkers rather than a preference model, the whole run is only as good as those checks, and writing them by hand is where the real thinking about the task happens.
The group baseline is the trick that makes GRPO cheap. Using the group mean in place of a learned value network removes a whole model from the loop, and watching it work made the method feel far less mysterious than the acronym suggests.
Verifiable rewards are what keep the model honest. Because a wrong answer scores zero no matter how confident the prose is, the model cannot bluff its way to a high reward. That single property is most of what makes math a good first domain for this kind of training.
References
- Shao, Z. et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning (GRPO). arXiv:2402.03300, 2024.
- Cobbe, K. et al. Training Verifiers to Solve Math Word Problems (GSM8K). arXiv:2110.14168, 2021.
- Ouyang, L. et al. Training language models to follow instructions with human feedback. NeurIPS, 2022.
- Hu, E. et al. LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685, 2021.
Code, model, and hackathon writeup.
- Code: gemma-3-1b-reasoning
- Model: gemma-3-1b-reasoning
- Writeup: Google Tunix Hackathon on Kaggle

