Gaze Attention, Qwen3-4B, 384px, LLaVA-OneVision data
Image model of the paper Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs (Junha Song, Byeongho Heo, Geonmo Gu, Jaegul Choo, Dongyoon Han, Sangdoo Yun; NAVER AI Lab, KAIST).
- Paper: https://arxiv.org/abs/2605.13080
- Code: https://github.com/junha1125/Gaze-Attention
- Project page: https://june-page.github.io/gaze-attention
Multimodal LLMs attend to all visual tokens at every generation step. With Gaze Attention, the cached keys of the visual tokens are grouped into spatial regions, and every text query attends only to the TopK regions that match it best, selected per layer, per head and per token. A few learnable context tokens per image or frame keep the global view of the scene. The KV cache stays complete, so the attended regions change from one generated token to the next.
Model
| Language model | Qwen3-4B-Instruct-2507 |
| Vision encoder | SigLIP2-So400m/16, 384px (finetuned) |
| Projector | 2-layer MLP |
| Visual tokens | 576 per image |
| Gaze regions | 16 regions of 6 x 6 tokens |
| Attended visual KV entries per query | 2 regions (72 tokens) + 4 context tokens |
| Context tokens | 4 per image |
| Parameters | 4.5B, stored in fp16 |
Usage
The model is loaded with the cambrian package of the Gaze Attention repository:
git clone https://github.com/junha1125/Gaze-Attention
cd Gaze-Attention
pip install torch==2.2.2 torchvision==0.17.2
pip install -e .
import copy
import torch
from PIL import Image
from cambrian.constants import DEFAULT_IMAGE_TOKEN, IMAGE_TOKEN_INDEX
from cambrian.conversation import conv_templates
from cambrian.mm_utils import process_images, tokenizer_image_token
from cambrian.model.builder import load_pretrained_model
tokenizer, model, image_processor, _ = load_pretrained_model("junha1125/gaze-attention-qwen3-4b-384-onevision", None, "cambrian_qwen3", device_map="cuda:0")
model.eval()
image = Image.open("assets/example2_dog.jpg").convert("RGB")
image_tensor = process_images([image], image_processor, model.config)
image_tensor = [x.to(dtype=torch.float16, device="cuda:0") for x in image_tensor]
conv = copy.deepcopy(conv_templates["qwen_3"])
conv.append_message(conv.roles[0], DEFAULT_IMAGE_TOKEN + "\nDescribe this image.")
conv.append_message(conv.roles[1], None)
input_ids = tokenizer_image_token(conv.get_prompt(), tokenizer, IMAGE_TOKEN_INDEX, return_tensors="pt").unsqueeze(0).to("cuda:0")
output_ids = model.generate(input_ids, images=image_tensor, image_sizes=[image.size], do_sample=False, max_new_tokens=512)
print(tokenizer.batch_decode(output_ids, skip_special_tokens=True)[0])
or python demo/image_inference.py --model-path junha1125/gaze-attention-qwen3-4b-384-onevision --image assets/example2_dog.jpg --question "Describe this image."
gaze_topk_ratio (0.1 in this checkpoint) is the fraction of gaze regions that a query attends to. It can be changed
when the model is loaded: load_pretrained_model(..., overwrite_config={"gaze_topk_ratio": 0.25}).
The generation config of Qwen3 samples by default; pass do_sample=False for reproducible outputs.
Training
| Stage | Trained parts | Data | Learning rate | Batch size |
|---|---|---|---|---|
| 1-1 | projector | Cambrian-Alignment (1.88M samples) | 1e-3 | 512 |
| 1-2 | projector, context tokens | Cambrian-Alignment | 1e-3 | 512 |
| 2 | vision encoder, projector, language model | LLaVA-OneVision single-image data (3.23M samples) + 40% of Cambrian-7M (2.83M samples) | 1e-5 (vision encoder 2e-6) | 256 |
Stage 2 uses the progressive TopK schedule: the fraction of attended regions decreases from 0.9 to 0.1 during the first 60% of
the training steps. Recipe: scripts/train/image_384 of the repository (3_finetune_gaze_onevision.sh).
This model is not part of the paper, which reports models trained on Cambrian-7M.
License
Apache License 2.0. The model is built on Qwen3-4B-Instruct-2507 and SigLIP2 (both Apache 2.0). The training datasets have their own licenses.
Citation
@inproceedings{song2026gazeattention,
title={Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal {LLM}s},
author={Junha Song and Byeongho Heo and Geonmo Gu and Jaegul Choo and Dongyoon Han and Sangdoo Yun},
booktitle={Third Conference on Language Modeling (CoLM)},
year={2026}
}
- Downloads last month
- 37
Model tree for junha1125/gaze-attention-qwen3-4b-384-onevision
Base model
Qwen/Qwen3-4B-Instruct-2507