Qwen3-1.7B-Base-KnK6-RL
Qwen3-1.7B-Base trained with verifiable-reward RL for 1000 steps on 6-person
Knights and Knaves logic puzzles.
Knights always tell the truth and knaves always lie. Given a set of statements, the model must assign a role to every character. Puzzles are generated by reasoning-gym; this model was trained on the 6-character configuration.
Results
Validation on 200 held-out KnK-6 puzzles, 64 samples per puzzle at temperature 0.6 / top_p 0.95:
| pass@1 | pass@64 | format | |
|---|---|---|---|
| Qwen3-1.7B-Base (step 0) | 0.015 | 0.323 | — |
| this model (step 1000) | 0.740 | 0.973 | 0.984 |
Coverage is not merely sharpened: pass@64 also rises from 0.323 to 0.973, and final pass@1 exceeds the starting pass@64 by more than a factor of two.
Prompt format
Puzzles are presented as a single user turn ending with the answer-format instruction, and the
answer is read from the last \boxed{}:
<puzzle text> Let's think step by step and output the final answer within \boxed{}.
The reward is 1 when the boxed assignment matches the unique solution and 0 otherwise. There is no partial credit and no format bonus.
Training
Trained with verl. Recipe:
| advantage estimator | maxrl |
| learning rate | 1e-6 |
| steps | 1000 |
| train batch size | 64 prompts |
| rollouts per prompt | 32 |
| rollout temperature | 1.0 |
| max response length | 4096 |
| KL penalty | none (use_kl_loss=False, use_kl_in_reward=False) |
| entropy bonus | 0 |
Intended use and limitations
This is a research artifact from a study of RL plasticity, not a general-purpose assistant. It
was optimized for a single puzzle family and its behaviour outside that distribution is not
characterized. It inherits the licence and limitations of Qwen/Qwen3-1.7B-Base.
- Downloads last month
- 337
Model tree for siddharthsingh18/Qwen3-1.7B-KnK6RL-Polaris-GRPO-step900-bestpass1
Base model
Qwen/Qwen3-1.7B-Base