kagakouko commited on
Commit
e732ce5
·
verified ·
1 Parent(s): b4fa224

Simplify model cards and add direct loading examples

Browse files
Files changed (1) hide show
  1. README.md +13 -28
README.md CHANGED
@@ -52,39 +52,24 @@ trajectories.
52
  <img src="https://raw.githubusercontent.com/ZJU-OmniAI/Spatial-Interactor/main/assets/readme/presentation-preview.webp?v=20260916" width="100%" alt="Spatial-Interactor 20-page presentation">
53
  </p>
54
 
55
- ## How Spatial-Interactor learns
56
-
57
- <p align="center">
58
- <img src="https://raw.githubusercontent.com/ZJU-OmniAI/Spatial-Interactor/main/assets/readme/opd.webp" width="100%" alt="On-Policy Distillation pipeline">
59
- </p>
60
-
61
- L1 and L2 establish local state-transition modeling. On L3, verifiable answer
62
- rewards supervise the result while same-prefix privileged distillation guides
63
- the intermediate reasoning process.
64
-
65
- ## Usage
66
-
67
- All four checkpoints are listed in the
68
- [Spatial-Interactor collection](https://huggingface.co/collections/kagakouko/spatial-interactor).
69
-
70
- Use the standard Transformers interface for the base model and load this
71
- repository in place of the base identifier:
72
 
73
  ```python
 
 
 
74
  model_id = "kagakouko/Spatial-Interactor-Qwen3-VL-8B"
 
 
 
 
75
  ```
76
 
77
- For video evaluation, preserve chronological frame order and use the frame
78
- budget specified by the target benchmark. The paper's main video evaluation
79
- uses 32 ordered frames.
80
-
81
- ## Training summary
82
-
83
- The SFT stage uses the reported L1-L2 split of LSI-108K together with the public
84
- spatial QA mixture described in the paper. OPD starts from that SFT checkpoint
85
- and combines verifiable answer rewards with CoT-only privileged
86
- self-distillation on long-horizon video questions. The visual encoder remains
87
- frozen while the language model and multimodal projector are updated.
88
 
89
  ## Citation
90
 
 
52
  <img src="https://raw.githubusercontent.com/ZJU-OmniAI/Spatial-Interactor/main/assets/readme/presentation-preview.webp?v=20260916" width="100%" alt="Spatial-Interactor 20-page presentation">
53
  </p>
54
 
55
+ ## Load
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
56
 
57
  ```python
58
+ import torch
59
+ from transformers import AutoModelForImageTextToText, AutoProcessor
60
+
61
  model_id = "kagakouko/Spatial-Interactor-Qwen3-VL-8B"
62
+ processor = AutoProcessor.from_pretrained(model_id)
63
+ model = AutoModelForImageTextToText.from_pretrained(
64
+ model_id, torch_dtype=torch.bfloat16, device_map="auto",
65
+ )
66
  ```
67
 
68
+ Use the base model's [image/video input format](https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct).
69
+ No privileged trace or additional teacher is needed for inference.
70
+ Weights, tokenizer, processor, and chat template are included.
71
+ See the [training guide](https://github.com/ZJU-OmniAI/Spatial-Interactor/blob/main/docs/TRAINING.md)
72
+ for SFT and OPD.
 
 
 
 
 
 
73
 
74
  ## Citation
75