--- license: apache-2.0 base_model: Qwen/Qwen3-VL-8B-Instruct library_name: transformers pipeline_tag: image-text-to-text tags: - vision-language - video - spatial-reasoning - embodied-ai - qwen ---

Spatial-Interactor

Spatial-Interactor Qwen3-VL-8B

Learning Spatial Reasoning through Interaction with the Observable Physical World

Project page Paper PDF Code Dataset

This is the full-parameter BF16 **Spatial-Interactor** checkpoint based on [Qwen/Qwen3-VL-8B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct). It learns local world-state and ego-motion transitions through supervised fine-tuning, then uses On-Policy Distillation (OPD) to integrate successive transitions over long trajectories. > The privileged transition trace is used only during training. At inference, > this checkpoint takes the same image/video and question inputs as its base > model, with no extra trace, reward model, or teacher branch. ## Overview

Spatial-Interactor overview: interaction trajectories, three-level curriculum, SFT and OPD, and spatial reasoning results

## Presentation ## Load ```python import torch from transformers import AutoModelForImageTextToText, AutoProcessor model_id = "kagakouko/Spatial-Interactor-Qwen3-VL-8B" processor = AutoProcessor.from_pretrained(model_id) model = AutoModelForImageTextToText.from_pretrained( model_id, torch_dtype=torch.bfloat16, device_map="auto", ) ``` Use the base model's [image/video input format](https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct). No privileged trace or additional teacher is needed for inference. Weights, tokenizer, processor, and chat template are included. See the [training guide](https://github.com/ZJU-OmniAI/Spatial-Interactor/blob/main/docs/TRAINING.md) for SFT and OPD. ## Citation For citation, use the [project BibTeX](https://zju-omniai.github.io/Spatial-Interactor/#citation). ## License This checkpoint is released under Apache-2.0, following the base model license. Users must also comply with licenses and terms governing input datasets and media.