ICML 2026 · Open weights & code

Strong vision-language reasoning, packed into INT4.

GRACE distills Qwen3-VL-8B into a 2B student and trains it for low-bit deployment from the start.

2B
student parameters, distilled from 8B
98%
of the GRACE BF16 benchmark average retained
INT4
real AWQ-packed language-model weights

Quality after compression

Average over HallusionBench, MMBench, ScienceQA, AI2D, MMMU, SEED-Bench, and MMStar using the released evaluation protocol.

ModelParametersFormatAverage
Qwen3-VL teacher8BBF1676.3
Qwen3-VL baseline2BBF1667.3
GRACE student2BBF1676.7
GRACE W4G1282BINT475.0
China Airlines aircraft used in the GRACE inference example

Example output

The model identifies the China Airlines livery and the Boeing 777-300ER, then grounds its description in the airport runway scene.

Generated by the released Qwen3-VL-2B-GRACE-W4G128-AWQ checkpoint. See the repository for the full output and settings.

Run the real packed checkpoint

The QAT repository is for research and repacking. Use the -AWQ repository below for genuine INT4 storage and kernels.

git clone https://github.com/ForeverBlue816/GRACE.git
cd GRACE
pip install torch==2.5.1 torchvision==0.20.1 \
  --index-url https://download.pytorch.org/whl/cu121
pip install -r requirements_inference.txt
pip install -e qwen-vl-utils/
python qwen-vl-finetune/scripts/deploy_awq_qwen.py \
  --load-packed ForeverBlue/Qwen3-VL-2B-GRACE-W4G128-AWQ \
  --image deployment/images/chinaairlines.jpg \
  --query "Describe this image in detail."