kagakouko's picture
Align the English introduction with the original film
71c5411 verified
|
Raw History Blame Contribute Delete
4.09 kB
---
license: other
license_name: qwen-research
license_link: https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE
base_model: Qwen/Qwen2.5-VL-3B-Instruct
library_name: transformers
pipeline_tag: image-text-to-text
tags:
- vision-language
- video
- spatial-reasoning
- embodied-ai
- qwen
---
<p align="center">
<img src="https://raw.githubusercontent.com/ZJU-OmniAI/Spatial-Interactor/main/assets/readme/icon.png" width="100" alt="Spatial-Interactor">
</p>
<h1 align="center">Spatial-Interactor Qwen2.5-VL-3B</h1>
<p align="center">
<strong>Learning Spatial Reasoning through Interaction with the Observable Physical World</strong>
</p>
<p align="center">
<a href="https://zju-omniai.github.io/Spatial-Interactor/"><img src="https://img.shields.io/badge/Project-Page-A56F59?style=flat-square&amp;labelColor=54534D" alt="Project page"></a>
<a href="https://zju-omniai.github.io/Spatial-Interactor/assets/paper.pdf?v=20260917"><img src="https://img.shields.io/badge/Paper-PDF-9B8255?style=flat-square&amp;labelColor=54534D" alt="Paper PDF"></a>
<a href="https://github.com/ZJU-OmniAI/Spatial-Interactor"><img src="https://img.shields.io/badge/Code-GitHub-738363?style=flat-square&amp;labelColor=54534D" alt="Code"></a>
<a href="https://huggingface.co/datasets/kagakouko/LSI-108K"><img src="https://img.shields.io/badge/LSI--108K-Dataset-887A9A?style=flat-square&amp;labelColor=54534D" alt="Dataset"></a>
</p>
This is the full-parameter BF16 **Spatial-Interactor** checkpoint based on
[Qwen/Qwen2.5-VL-3B-Instruct](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct). It learns local
world-state and ego-motion transitions through supervised fine-tuning, then
uses On-Policy Distillation (OPD) to integrate successive transitions over long
trajectories.
> The privileged transition trace is used only during training. At inference,
> this checkpoint takes the same image/video and question inputs as its base
> model, with no extra trace, reward model, or teacher branch.
## Overview
<p align="center">
<img src="https://raw.githubusercontent.com/ZJU-OmniAI/Spatial-Interactor/main/assets/readme/overview.webp?v=20260917" width="100%" alt="Spatial-Interactor overview: interaction trajectories, three-level curriculum, SFT and OPD, and spatial reasoning results">
</p>
<video src="https://huggingface.co/kagakouko/Spatial-Interactor-Qwen2.5-VL-3B/resolve/main/assets/presentation/spatial-interactor-intro-en.mp4?v=20260924-faithful" controls autoplay muted loop playsinline preload="metadata" poster="https://huggingface.co/kagakouko/Spatial-Interactor-Qwen2.5-VL-3B/resolve/main/assets/presentation/spatial-interactor-intro-en-poster.webp?v=20260924-faithful" width="100%"></video>
## Presentation
<video src="https://huggingface.co/kagakouko/Spatial-Interactor-Qwen2.5-VL-3B/resolve/main/assets/presentation/spatial-interactor-presentation-en.mp4?v=20260924" controls autoplay muted loop playsinline preload="metadata" poster="https://huggingface.co/kagakouko/Spatial-Interactor-Qwen2.5-VL-3B/resolve/main/assets/presentation/presentation-en-poster.webp?v=20260924" width="100%"></video>
## Load
```python
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
model_id = "kagakouko/Spatial-Interactor-Qwen2.5-VL-3B"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id, torch_dtype=torch.bfloat16, device_map="auto",
)
```
Use the base model's [image/video input format](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct).
No privileged trace or additional teacher is needed for inference.
Weights, tokenizer, processor, and chat template are included.
See the [training guide](https://github.com/ZJU-OmniAI/Spatial-Interactor/blob/main/docs/TRAINING.md)
for SFT and OPD.
## Citation
For citation, use the [project BibTeX](https://zju-omniai.github.io/Spatial-Interactor/#citation).
## License
This checkpoint follows the Qwen Research License of the base model. Users must also comply with licenses and terms governing
input datasets and media.