openai/gsm8k
Benchmark • Updated • 17.6k • 1.24M • 1.59k
GRPO experiment from TinkerRL-Bench world-class experiment suite.
[
1.0,
0.875,
1.0,
1.0,
0.75,
1.0,
0.75,
1.0,
0.75,
0.875,
1.0,
1.0,
1.0,
0.5,
0.875,
0.375,
0.875,
0.75,
0.875,
0.75
]
@misc{tinker-rl-bench-2026,
title={TinkerRL-Bench: A Unified Benchmark for RL Post-Training},
author={Arvind C R and Sandhya Jeyaraj and Madhu Kumara L and Mohammad Rafi and Dhruva N Murthy and Arumugam K},
year={2026},
url={https://github.com/arvindcr4/tinker-rl-lab}
}
Base model
moonshotai/Kimi-K2-Thinking