--- license: apache-2.0 base_model: Qwen/Qwen3-VL-4B-Instruct library_name: transformers pipeline_tag: image-text-to-text tags: - spatial-reasoning - vision-language-model - reinforcement-learning - self-distillation --- # GPD-4B [![arXiv](https://img.shields.io/badge/arXiv-GPD-b31b1b?logo=arxiv&logoColor=white)](https://arxiv.org/abs/2610.12355) [![Code](https://img.shields.io/badge/Code-GPD-181717?logo=github&logoColor=white)](https://github.com/ZJU-REAL/GPD) [![Model](https://img.shields.io/badge/%F0%9F%A4%97%20Model-GPD--2B-ffc107)](https://huggingface.co/xinyili0624/GPD-2B) [![Data](https://img.shields.io/badge/%F0%9F%A4%97%20Data-GPD--15k-ffc107)](https://huggingface.co/datasets/xinyili0624/GPD-15k) Official checkpoint of **GPD** (*Geometry-Privileged Distillation*) from the paper [**Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models**](https://huggingface.co/papers/2610.12355). - Base model: [Qwen3-VL-4B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct) - Training data: Mixed-15k (VSI 10k + SPAR 4k + MindCube 1k), `text_routed` variant - Training: GRPO + privileged KL on incorrect trajectories (`kl_coef=0.003`), 111 steps - Inference: RGB-only, same interface as the base model ## Results | Model | VSI-Bench | Avg. (MindCube / SPAR / MMSI / ViewSpatial) | | ----- | --------- | ------------------------------------------- | | GRPO | 56.4 | 35.7 | | OPSD (answer privilege) | 55.7 | 31.5 | | **GPD** | **57.1** | **37.6** | ## Usage ```python from transformers import AutoProcessor, Qwen3VLForConditionalGeneration model_id = "xinyili0624/GPD-4B" model = Qwen3VLForConditionalGeneration.from_pretrained(model_id, torch_dtype="auto", device_map="auto") processor = AutoProcessor.from_pretrained(model_id) ``` Usage is identical to Qwen3-VL-4B-Instruct; see the base model card for image / video inference examples.