--- library_name: transformers base_model: Qwen/Qwen3.5-9B pipeline_tag: text-generation license: apache-2.0 tags: - qwen3.5 - supervised-fine-tuning - webshop - external-guidance - workflow --- # Qwen3.5-9B-OutsideTheBox-SFT This repository provides the WebShop supervised fine-tuning checkpoint developed for the paper **Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?** ![Thinking Outside the Box](thinkoutsidethebox.png) ## Model Description The model is initialized from [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) and fine-tuned to act in the WebShop text environment under external workflow guidance. Each training example contains the WebShop system prompt, the shopping instruction and interaction history with an injected misleading workflow, and the correct next action from a successful trajectory; only the assistant action is included in the loss. ## Training Data The training set is constructed exclusively from the WebShop `TRAIN` split and contains 2,023 tasks and 8,203 turn-level examples. Official WebShop test tasks are not used for training. ## Training Configuration | Setting | Value | |---|---:| | Base model | `Qwen/Qwen3.5-9B` | | Training | Full-parameter SFT | | Workflow condition | Misleading workflow | | Epochs | 2 | | Learning rate | `1e-6` | | Global batch size | 32 | | Maximum sequence length | 12,288 | | Precision | BF16 | ## Intended Use This checkpoint is intended for research on WebShop agents, external workflow utilization, and robustness to misleading guidance. Reproducing the reported results requires the WebShop environment and the interaction harness used in the paper. ## Loading the Model ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "ElvisWang111/Qwen3.5-9B-OutsideTheBox-SFT" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype="auto", device_map="auto", ) ```