FineEnvs/SmolDataEnvs
RL Environment • Updated • 5.39k • 2.46k • 50
5.5K+ RL tasks for hill-climbing small models in code and data science. Deterministic grading, no LLM judge.
GRPO training logs for Qwen3.5 on SmolDataEnvs
Note Trackio logs for the GRPO runs behind the curves on the dataset cards.