Qwen3-14B τ-bench GRPO checkpoints

LoRA adapters (r=16) from a GRPO reinforcement-learning run on unsloth/Qwen3-14B-unsloth-bnb-4bit against τ-bench's retail domain, using τ-bench's own calculate_reward() as the RL reward (no custom scoring). Full write-up, code, and results: qwen3-14b-tau-train.

Checkpoints

Folder Training round Eval probe (20 fixed real test tasks, temp=0)
round-0/ 1 (12/16 rollouts contributed gradient) 0.55 (0.50 on independent rerun)
round-5/ 5 0.35
round-10/ 10 0.30 (0.35 on rerun)

round-0 is the best-performing of the three — more training made the eval probe score go down, a decline confirmed on an independent rerun with full dialogues saved (see the GitHub repo's results/probe_transcripts/). This is a documented negative result: this specific GRPO configuration (batch size 4 tasks/round, lr=1e-5, 12 rounds) did not improve the model over its lightly-trained starting point.

Augmentation SFT checkpoints (aug-*)

Fresh LoRA (r=16) trained from the base model on teacher conversations rebuilt so that the agent opens every order and confirms every change before a database change (511 steps, 1 epoch, lr 5e-5). aug-step-96/192/288/384/480 are intermediate checkpoints, aug-epoch-1 is the final adapter (also published as teacher57/qwen3-14b-tau-lookup-confirm-augmented), aug-train_log.jsonl is the loss log. Results (not statistically significant): τ-bench retail pass^1 44.6% (starting adapter 41.7%, distilled 45.7%); τ³ retail default settings pass^1 47.8% vs 43.4%, pass^2 35.1% vs 25.4%. Details: GitHub branch augmentation-experiment.

Training loss of the augmented SFT

pass^k on τ-bench and τ³: starting adapter, distilled, augmented

Causes of failed τ-bench rollouts, three models

Usage

Each folder is a standard PEFT LoRA adapter directory (adapter_config.json + adapter_model.safetensors), loadable against the base model either via peft/transformers, or served from vLLM with --enable-lora and --lora-modules <name>=<path>.

Combo-pool GRPO run

A second GRPO run, started from the same round-0 adapter, trained only on the 56 combo train tasks the model solved some of the time (lr 1e-5, 4 tasks x 4 rollouts per round). Probe: the same 20 held-out test tasks at temperature 0.

Folder What it is Eval probe
(the starting adapter, round-0/) baseline for this run 0.40
combo-round-0/ after the first update of the combo run 0.45
combo-round-5/ after 5 rounds 0.20

The run was stopped automatically in round 7 by the decline rule. The probe has 20 tasks and is noisy (about +/-0.1; the same starting adapter scored 0.55 and 0.50 in the earlier setup), so treat the gap as suggestive. Details, per-task results and transcripts: the GitHub repo (section 8 of the README).

Gentle combo-pool GRPO run

The same run as above with a 10x lower learning rate (1e-6), same start (the round-0 adapter) and same 56 combo tasks. Checkpoints were saved and probed (20 held-out test tasks, temperature 0) at rounds 1, 3, 5 and 10.

Folder Eval probe
(the starting adapter, round-0/) 0.40
gentle-round-1/ 0.50
gentle-round-3/ 0.45
gentle-round-5/ 0.60
gentle-round-10/ 0.40

None of these is measurably better than the starting adapter. A full 115-task test with two trials per task at temperature 0.7 gave gentle-round-5 pass^1 40.4% and pass^2 25.2%, against 37.8% and 21.7% for the starting adapter (paired pass^2 p = 0.57). The probe has only 20 tasks and is noisy (about +/-0.1), so the 0.60 at round 5 was most likely chance. Details, per-task results and the full write-up: the GitHub repo (README sections 8 and 9).

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for teacher57/qwen3-14b-tau-grpo-checkpoints

Finetuned
Qwen/Qwen3-14B
Adapter
(31)
this model