Qwen3-8B-UserRL-Agent

The UserRL baseline from MIMESIS: Qwen3-8B trained with multi-turn GRPO, with GPT-5.5 as the simulated user, following UserRL. The paper reports it as "GRPO w. GPT-5.5" in the main text and as UserRL in Appendix D.

Paper · Code · Project page · Collection

Results

Mean task score across eight Gym environments, three of them held out from training, under nine evaluation user models that were not used in training:

Training condition Mean score
GRPO w. GPT-5.5 (this model) 26.10
GRPO w. MIMESIS-9B 29.54

With the GRPO objective fixed, replacing GPT-5.5 with MIMESIS-9B as the training user improves the mean score under every evaluation user. Per-environment and per-user scores are in Appendix D of the paper.

Usage

The agent acts in the UserRL Gym environments through a tool call. Training and evaluation scripts, including evaluation under any user model, are in the agent directory of the code release.

Citation

@article{phan2026mimesis,
  title   = {{MIMESIS}: Learning User Simulators as Training Environments for Interactive Agents},
  author  = {Phan, Hoang and Huynh, Dat and Zhmoginov, Andrey and Zeng, Qi and Mu, Wancen and Cao, Yue and Bi, Shengjie and He, Yun and Oh, Changdae and Lei, Deren},
  journal = {arXiv preprint arXiv:2610.09484},
  year    = {2026}
}
Downloads last month
254
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for phanviethoang1512/Qwen3-8B-UserRL-Agent

Finetuned
Qwen/Qwen3-8B
Finetuned
(2174)
this model

Collection including phanviethoang1512/Qwen3-8B-UserRL-Agent

Paper for phanviethoang1512/Qwen3-8B-UserRL-Agent