Gaze Attention, Qwen3-4B, 384px, LLaVA-OneVision data

Image model of the paper Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal LLMs (Junha Song, Byeongho Heo, Geonmo Gu, Jaegul Choo, Dongyoon Han, Sangdoo Yun; NAVER AI Lab, KAIST).

Multimodal LLMs attend to all visual tokens at every generation step. With Gaze Attention, the cached keys of the visual tokens are grouped into spatial regions, and every text query attends only to the TopK regions that match it best, selected per layer, per head and per token. A few learnable context tokens per image or frame keep the global view of the scene. The KV cache stays complete, so the attended regions change from one generated token to the next.

Model

Language model Qwen3-4B-Instruct-2507
Vision encoder SigLIP2-So400m/16, 384px (finetuned)
Projector 2-layer MLP
Visual tokens 576 per image
Gaze regions 16 regions of 6 x 6 tokens
Attended visual KV entries per query 2 regions (72 tokens) + 4 context tokens
Context tokens 4 per image
Parameters 4.5B, stored in fp16

Usage

The model is loaded with the cambrian package of the Gaze Attention repository:

git clone https://github.com/junha1125/Gaze-Attention
cd Gaze-Attention
pip install torch==2.2.2 torchvision==0.17.2
pip install -e .
import copy
import torch
from PIL import Image

from cambrian.constants import DEFAULT_IMAGE_TOKEN, IMAGE_TOKEN_INDEX
from cambrian.conversation import conv_templates
from cambrian.mm_utils import process_images, tokenizer_image_token
from cambrian.model.builder import load_pretrained_model

tokenizer, model, image_processor, _ = load_pretrained_model("junha1125/gaze-attention-qwen3-4b-384-onevision", None, "cambrian_qwen3", device_map="cuda:0")
model.eval()

image = Image.open("assets/example2_dog.jpg").convert("RGB")
image_tensor = process_images([image], image_processor, model.config)
image_tensor = [x.to(dtype=torch.float16, device="cuda:0") for x in image_tensor]

conv = copy.deepcopy(conv_templates["qwen_3"])
conv.append_message(conv.roles[0], DEFAULT_IMAGE_TOKEN + "\nDescribe this image.")
conv.append_message(conv.roles[1], None)
input_ids = tokenizer_image_token(conv.get_prompt(), tokenizer, IMAGE_TOKEN_INDEX, return_tensors="pt").unsqueeze(0).to("cuda:0")

output_ids = model.generate(input_ids, images=image_tensor, image_sizes=[image.size], do_sample=False, max_new_tokens=512)
print(tokenizer.batch_decode(output_ids, skip_special_tokens=True)[0])

or python demo/image_inference.py --model-path junha1125/gaze-attention-qwen3-4b-384-onevision --image assets/example2_dog.jpg --question "Describe this image."

gaze_topk_ratio (0.1 in this checkpoint) is the fraction of gaze regions that a query attends to. It can be changed when the model is loaded: load_pretrained_model(..., overwrite_config={"gaze_topk_ratio": 0.25}). The generation config of Qwen3 samples by default; pass do_sample=False for reproducible outputs.

Training

Stage Trained parts Data Learning rate Batch size
1-1 projector Cambrian-Alignment (1.88M samples) 1e-3 512
1-2 projector, context tokens Cambrian-Alignment 1e-3 512
2 vision encoder, projector, language model LLaVA-OneVision single-image data (3.23M samples) + 40% of Cambrian-7M (2.83M samples) 1e-5 (vision encoder 2e-6) 256

Stage 2 uses the progressive TopK schedule: the fraction of attended regions decreases from 0.9 to 0.1 during the first 60% of the training steps. Recipe: scripts/train/image_384 of the repository (3_finetune_gaze_onevision.sh). This model is not part of the paper, which reports models trained on Cambrian-7M.

License

Apache License 2.0. The model is built on Qwen3-4B-Instruct-2507 and SigLIP2 (both Apache 2.0). The training datasets have their own licenses.

Citation

@inproceedings{song2026gazeattention,
  title={Gaze Attention: Query-Adaptive Visual Routing for Efficient Multimodal {LLM}s},
  author={Junha Song and Byeongho Heo and Geonmo Gu and Jaegul Choo and Dongyoon Han and Sangdoo Yun},
  booktitle={Third Conference on Language Modeling (CoLM)},
  year={2026}
}
Downloads last month
37
Safetensors
Model size
5B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for junha1125/gaze-attention-qwen3-4b-384-onevision

Finetuned
(2381)
this model

Datasets used to train junha1125/gaze-attention-qwen3-4b-384-onevision

Paper for junha1125/gaze-attention-qwen3-4b-384-onevision