---
license: other
license_name: qwen-research
license_link: https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE
base_model: Qwen/Qwen2.5-VL-3B-Instruct
library_name: transformers
pipeline_tag: image-text-to-text
tags:
- vision-language
- video
- spatial-reasoning
- embodied-ai
- qwen
---
Spatial-Interactor Qwen2.5-VL-3B
Learning Spatial Reasoning through Interaction with the Observable Physical World
This is the full-parameter BF16 **Spatial-Interactor** checkpoint based on
[Qwen/Qwen2.5-VL-3B-Instruct](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct). It learns local
world-state and ego-motion transitions through supervised fine-tuning, then
uses On-Policy Distillation (OPD) to integrate successive transitions over long
trajectories.
> The privileged transition trace is used only during training. At inference,
> this checkpoint takes the same image/video and question inputs as its base
> model, with no extra trace, reward model, or teacher branch.
## Overview
## Presentation
## Load
```python
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
model_id = "kagakouko/Spatial-Interactor-Qwen2.5-VL-3B"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id, torch_dtype=torch.bfloat16, device_map="auto",
)
```
Use the base model's [image/video input format](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct).
No privileged trace or additional teacher is needed for inference.
Weights, tokenizer, processor, and chat template are included.
See the [training guide](https://github.com/ZJU-OmniAI/Spatial-Interactor/blob/main/docs/TRAINING.md)
for SFT and OPD.
## Citation
For citation, use the [project BibTeX](https://zju-omniai.github.io/Spatial-Interactor/#citation).
## License
This checkpoint follows the Qwen Research License of the base model. Users must also comply with licenses and terms governing
input datasets and media.