Instructions to use teacher57/qwen3-14b-tau-lookup-confirm-augmented with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use teacher57/qwen3-14b-tau-lookup-confirm-augmented with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/Qwen3-14B-unsloth-bnb-4bit") model = PeftModel.from_pretrained(base_model, "teacher57/qwen3-14b-tau-lookup-confirm-augmented") - Notebooks
- Google Colab
- Kaggle
Qwen3-14B LoRA trained to open every order and confirm every change (τ-bench retail)
A QLoRA adapter (r=16) for unsloth/Qwen3-14B-unsloth-bnb-4bit, trained from the base on the 180 Qwen3-32B teacher conversations of the earlier distillation experiment, rebuilt by rule so that the agent (1) loads the account and opens every order of the customer before anything else, and (2) lists exactly what it is about to change and gets an explicit "yes" before every database change. 178 conversations (2 dropped for changing an order before authenticating), 710 lookups and 309 confirmations inserted from the retail database, 1 epoch = 511 steps at batch 8, lr 5e-5, loss 0.65 → 0.30. Code and write-up: GitHub, section 11. Data and all conversations: teacher57/tau-retail-distillation (folders augmentation/, retail_eval_conversations/, tau3_eval/). Intermediate checkpoints (steps 96 to 480) are in teacher57/qwen3-14b-tau-grpo-checkpoints as aug-step-*.
Read this first: the improvement over the baselines is small and not statistically significant. τ-bench retail (115 tasks, 4 trials, temperature 0.7, GPT-4o customer): 44.6% pass^1 vs 41.7% for the starting adapter and 45.7% for the distilled model (paired: +2.8 vs control, p = 0.37; −1.1 vs distilled, p = 0.72). τ³ retail default settings (114 tasks, 2 trials, temperature 0, GPT-4.1 customer): 47.8% vs 43.4% for the starting adapter (+4.4 points, p = 0.33); pass^2 35.1% vs 25.4% (+9.6, p = 0.063).
Did it learn the habits?
| starting adapter | distilled | this adapter | |
|---|---|---|---|
| changes that came right after an explicit yes (τ-bench) | 44% | 37% | 59% |
| tasks with 2+ orders where all orders were opened before the first change | 18% | 21% | 24% |
| opened an order straight after loading the account | 66% | 59% | 37% |
Confirmation: yes. Lookup: no. The "open every order" decision was only 178 examples (about 1% of the agent text), against about 3,900 unchanged ones, and the model keeps reasoning "which order does the user mean, I should ask". We built a habit-only dataset (only lookup and confirmation turns carry loss, 898 conversations) to fix that; it has not been trained yet.
Failure causes (rule-based classifier, checked against about 20 hand-read dialogs; failed τ-bench rollouts)
Lookup (72) is unchanged against the distilled model (75); changing before a yes fell (52 vs 68); wrong or invented values rose (72 vs 52). The model also takes more steps: 27 of 460 τ-bench rollouts hit the 25-step limit (control 7, distilled 10) and 19 of 228 τ³ rollouts hit a 20-minute limit (baseline 1).
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = "unsloth/Qwen3-14B-unsloth-bnb-4bit"
model = AutoModelForCausalLM.from_pretrained(base, device_map="auto")
model = PeftModel.from_pretrained(model, "teacher57/qwen3-14b-tau-lookup-confirm-augmented")
tok = AutoTokenizer.from_pretrained("teacher57/qwen3-14b-tau-lookup-confirm-augmented")
For evaluation we served it with vLLM 0.11.0 (--enable-lora --max-loras 3 --max-lora-rank 16, tool parser hermes, reasoning parser qwen3).
Limitations
Retail domain only; trained and tested with GPT-4o or GPT-4.1 as the simulated customer; one training run and 2 to 4 trials per task, so differences of a few points are within noise; the confirmation text in training is built from the database and always correct, while at test time the model writes its own and still picks wrong items often.
- Downloads last month
- 20
Model tree for teacher57/qwen3-14b-tau-lookup-confirm-augmented
Base model
Qwen/Qwen3-14B-Base

