R2 — Outcome contrast (DPO)

Part of a five-regime developmental sweep of post-training methods for dialogue-game competence (LM Playschool Challenge 2026, team DAIR).

DPO on top of the merged R1 model, using 3,418 first-move preference pairs built from playpen-data: for a fixed (game, experiment, task id, role), chosen is the first move of a successful episode and rejected the first move of an aborted one (falling back to lost). Abort-prone games oversampled 4x. One epoch, LoRA r=16/alpha=32, lr 5e-6, beta=0.1; frozen R1 as reference.

Effect: 55.61 -> 67.39. Unlike R1, the gain is in move quality (%-played stays flat), and it exceeds the untuned Qwen3.5-27B reference on the organisers' leaderboard. This checkpoint was our challenge submission.

All numbers are clemscore / statscore on the playpen validation split, measured in a single frozen environment (Python 3.11, clemcore pinned via playpen, clembench pinned requirements) with two upstream fixes applied: a division-by-zero guard in the privateshared Game Master and the punkt_tab NLTK resource for the IFEval scorer. Earlier revisions of this card reported numbers from an unpinned environment; see the paper for the environment-sensitivity analysis.

Checkpoint family (LM Playschool challenge, team DAIR)

Regime Repo clem stat
R1 imitation (SFT) lm-playschool-qwen3.5-2b-sft 55.61 43.87
R2 outcome contrast (DPO) lm-playschool-qwen3.5-2b-sft-dpo 67.39 44.72
R3 self-imitation (SFT) lm-playschool-qwen3.5-2b-iter3 61.06 44.01
R4 corrective feedback (DPO) lm-playschool-qwen3.5-2b-iter4 67.64 44.31
R5 GRPO (control) lm-playschool-qwen3.5-2b-grpo-base-s42 62.43 44.19
R5 GRPO + RND lm-playschool-qwen3.5-2b-grpo-rnd-s42 67.44 43.53

Base model: Qwen3.5-2B (13.63 / 44.22 in the same environment). Paper: Raising a Small Language Model: From Imitation to Curiosity in Dialogue Games (LM Playschool Challenge 2026).

Downloads last month
404
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for varadsrivastava/lm-playschool-qwen3.5-2b-sft-dpo

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(348)
this model

Collection including varadsrivastava/lm-playschool-qwen3.5-2b-sft-dpo