Qwen3.5-9B τ²-bench OPD LoRA (GKD)

PEFT adapter from on-policy distillation — base weights stay on Qwen/Qwen3.5-9B.

Field Value
TRAIN_MODE exp_fast
LoRA rank 32
W&B https://wandb.ai/alchemxz/decagon-posttraining-opd/runs/cibzral8

vLLM (runtime merge)

vllm serve Qwen/Qwen3.5-9B --enable-lora --lora-modules opd=lilyzhng/qwen3.5-9b-tau2-opd-lora \
  --max-lora-rank 32 --dtype bfloat16
Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lilyzhng/qwen3.5-9b-tau2-opd-lora

Finetuned
Qwen/Qwen3.5-9B
Adapter
(598)
this model