Qwen3-14B τ-bench GRPO checkpoints
LoRA adapters (r=16) from a GRPO reinforcement-learning run on
unsloth/Qwen3-14B-unsloth-bnb-4bit against τ-bench's
retail domain, using τ-bench's own calculate_reward() as the RL reward (no custom
scoring). Full write-up, code, and results:
qwen3-14b-tau-train.
Checkpoints
| Folder | Training round | Eval probe (20 fixed real test tasks, temp=0) |
|---|---|---|
round-0/ |
1 (12/16 rollouts contributed gradient) | 0.55 (0.50 on independent rerun) |
round-5/ |
5 | 0.35 |
round-10/ |
10 | 0.30 (0.35 on rerun) |
round-0 is the best-performing of the three — more training made the eval
probe score go down, a decline confirmed on an independent rerun with full
dialogues saved (see the GitHub repo's results/probe_transcripts/). This is a
documented negative result: this specific GRPO configuration (batch size 4
tasks/round, lr=1e-5, 12 rounds) did not improve the model over its
lightly-trained starting point.
Augmentation SFT checkpoints (aug-*)
Fresh LoRA (r=16) trained from the base model on teacher conversations rebuilt so that the agent opens every order and confirms every change before a database change (511 steps, 1 epoch, lr 5e-5). aug-step-96/192/288/384/480 are intermediate checkpoints, aug-epoch-1 is the final adapter (also published as teacher57/qwen3-14b-tau-lookup-confirm-augmented), aug-train_log.jsonl is the loss log. Results (not statistically significant): τ-bench retail pass^1 44.6% (starting adapter 41.7%, distilled 45.7%); τ³ retail default settings pass^1 47.8% vs 43.4%, pass^2 35.1% vs 25.4%. Details: GitHub branch augmentation-experiment.
Usage
Each folder is a standard PEFT LoRA adapter directory (adapter_config.json +
adapter_model.safetensors), loadable against the base model either via
peft/transformers, or served from vLLM with --enable-lora and
--lora-modules <name>=<path>.
Combo-pool GRPO run
A second GRPO run, started from the same round-0 adapter, trained only on the 56 combo train tasks the model solved some of the time (lr 1e-5, 4 tasks x 4 rollouts per round). Probe: the same 20 held-out test tasks at temperature 0.
| Folder | What it is | Eval probe |
|---|---|---|
(the starting adapter, round-0/) |
baseline for this run | 0.40 |
combo-round-0/ |
after the first update of the combo run | 0.45 |
combo-round-5/ |
after 5 rounds | 0.20 |
The run was stopped automatically in round 7 by the decline rule. The probe has 20 tasks and is noisy (about +/-0.1; the same starting adapter scored 0.55 and 0.50 in the earlier setup), so treat the gap as suggestive. Details, per-task results and transcripts: the GitHub repo (section 8 of the README).
Gentle combo-pool GRPO run
The same run as above with a 10x lower learning rate (1e-6), same start (the round-0 adapter) and same 56 combo tasks. Checkpoints were saved and probed (20 held-out test tasks, temperature 0) at rounds 1, 3, 5 and 10.
| Folder | Eval probe |
|---|---|
(the starting adapter, round-0/) |
0.40 |
gentle-round-1/ |
0.50 |
gentle-round-3/ |
0.45 |
gentle-round-5/ |
0.60 |
gentle-round-10/ |
0.40 |
None of these is measurably better than the starting adapter. A full 115-task test with two trials per task at
temperature 0.7 gave gentle-round-5 pass^1 40.4% and pass^2 25.2%, against 37.8% and 21.7% for the starting adapter
(paired pass^2 p = 0.57). The probe has only 20 tasks and is noisy (about +/-0.1), so the 0.60 at round 5 was most
likely chance. Details, per-task results and the full write-up: the GitHub repo (README sections 8 and 9).
Model tree for teacher57/qwen3-14b-tau-grpo-checkpoints
Base model
Qwen/Qwen3-14B-Base

