--- license: other license_name: qwen-research license_link: https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE base_model: Qwen/Qwen2.5-VL-3B-Instruct library_name: transformers pipeline_tag: image-text-to-text tags: - vision-language - video - spatial-reasoning - embodied-ai - qwen ---

Spatial-Interactor

Spatial-Interactor Qwen2.5-VL-3B

Learning Spatial Reasoning through Interaction with the Observable Physical World

Project page Paper PDF Code Dataset

This is the full-parameter BF16 **Spatial-Interactor** checkpoint based on [Qwen/Qwen2.5-VL-3B-Instruct](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct). It learns local world-state and ego-motion transitions through supervised fine-tuning, then uses On-Policy Distillation (OPD) to integrate successive transitions over long trajectories. > The privileged transition trace is used only during training. At inference, > this checkpoint takes the same image/video and question inputs as its base > model, with no extra trace, reward model, or teacher branch. ## Overview

Spatial-Interactor overview: interaction trajectories, three-level curriculum, SFT and OPD, and spatial reasoning results

## Presentation ## Load ```python import torch from transformers import AutoModelForImageTextToText, AutoProcessor model_id = "kagakouko/Spatial-Interactor-Qwen2.5-VL-3B" processor = AutoProcessor.from_pretrained(model_id) model = AutoModelForImageTextToText.from_pretrained( model_id, torch_dtype=torch.bfloat16, device_map="auto", ) ``` Use the base model's [image/video input format](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct). No privileged trace or additional teacher is needed for inference. Weights, tokenizer, processor, and chat template are included. See the [training guide](https://github.com/ZJU-OmniAI/Spatial-Interactor/blob/main/docs/TRAINING.md) for SFT and OPD. ## Citation For citation, use the [project BibTeX](https://zju-omniai.github.io/Spatial-Interactor/#citation). ## License This checkpoint follows the Qwen Research License of the base model. Users must also comply with licenses and terms governing input datasets and media.