Qwen3-1.7B-Base-KnK6-RL

Qwen3-1.7B-Base trained with verifiable-reward RL for 1000 steps on 6-person Knights and Knaves logic puzzles.

Knights always tell the truth and knaves always lie. Given a set of statements, the model must assign a role to every character. Puzzles are generated by reasoning-gym; this model was trained on the 6-character configuration.

Results

Validation on 200 held-out KnK-6 puzzles, 64 samples per puzzle at temperature 0.6 / top_p 0.95:

pass@1 pass@64 format
Qwen3-1.7B-Base (step 0) 0.015 0.323 —
this model (step 1000) 0.740 0.973 0.984

Coverage is not merely sharpened: pass@64 also rises from 0.323 to 0.973, and final pass@1 exceeds the starting pass@64 by more than a factor of two.

Prompt format

Puzzles are presented as a single user turn ending with the answer-format instruction, and the answer is read from the last \boxed{}:

<puzzle text> Let's think step by step and output the final answer within \boxed{}.

The reward is 1 when the boxed assignment matches the unique solution and 0 otherwise. There is no partial credit and no format bonus.

Training

Trained with verl. Recipe:

advantage estimator maxrl
learning rate 1e-6
steps 1000
train batch size 64 prompts
rollouts per prompt 32
rollout temperature 1.0
max response length 4096
KL penalty none (use_kl_loss=False, use_kl_in_reward=False)
entropy bonus 0

Intended use and limitations

This is a research artifact from a study of RL plasticity, not a general-purpose assistant. It was optimized for a single puzzle family and its behaviour outside that distribution is not characterized. It inherits the licence and limitations of Qwen/Qwen3-1.7B-Base.

Downloads last month
337
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for siddharthsingh18/Qwen3-1.7B-KnK6RL-Polaris-GRPO-step900-bestpass1

Finetuned
(441)
this model