How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="ElvisWang111/Qwen3.5-9B-OutsideTheBox-RL")
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("ElvisWang111/Qwen3.5-9B-OutsideTheBox-RL")
model = AutoModelForCausalLM.from_pretrained("ElvisWang111/Qwen3.5-9B-OutsideTheBox-RL", device_map="auto")
messages = [
    {"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))
Quick Links

Qwen3.5-9B-OutsideTheBox-RL

This repository provides the WebShop environment-reward checkpoint developed for the paper Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

Thinking Outside the Box

Model Description

The model is initialized from Qwen3.5-9B-OutsideTheBox-SFT and further optimized through multi-turn interaction with the WebShop environment. Training uses task instructions paired with grounded good workflows and directly rewards the final raw WebShop outcome in [0, 1]; it does not train on bad-workflow prompts.

RL Data Selection

Candidate tasks come exclusively from successful trajectories in the WebShop TRAIN split. The SFT checkpoint is sampled eight times per task, and tasks with non-constant raw rewards are retained. This produces 153 selected tasks, deterministically divided into 139 training tasks and a disjoint 14-task probe set; official WebShop test tasks are not used for selection or training.

Training Configuration

Setting Value
Initial policy Qwen3.5-9B-OutsideTheBox-SFT
Training Multi-turn environment GRPO
Workflow condition Grounded good workflow
Reward Raw WebShop task outcome [0, 1]
Generations per prompt 8
Epochs 1
Learning rate 1e-6
KL coefficient 0
Loss DAPO
Maximum tool calls 10
Maximum completion length 4,096 tokens

Intended Use

This checkpoint is intended for research on WebShop agents, outcome-based reinforcement learning, and selective reliance on external workflows. Reproducing the reported results requires the WebShop environment and the multi-turn interaction harness used in the paper.

Loading the Model

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "ElvisWang111/Qwen3.5-9B-OutsideTheBox-RL"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)
Downloads last month
247
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ElvisWang111/Qwen3.5-9B-OutsideTheBox-RL

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(1)
this model

Collection including ElvisWang111/Qwen3.5-9B-OutsideTheBox-RL