--- title: GRACE-VLM emoji: 🦢 colorFrom: blue colorTo: purple sdk: static app_file: index.html pinned: true short_description: Deployable INT4 Qwen3-VL through GRACE distillation. models: - ForeverBlue/Qwen3-VL-2B-GRACE-W4G128-AWQ - ForeverBlue/Qwen3-VL-2B-GRACE-BF16 tags: - vision-language-model - multimodal - int4 - awq - knowledge-distillation - arxiv:2601.22709 --- # GRACE-VLM Project showcase for **GRACE-VLM: INT4 Quantization-Aware Distillation for Vision-Language Models**, accepted at ICML 2026. Read the paper at [arXiv:2601.22709](https://arxiv.org/abs/2601.22709). This free static Space presents the benchmark results, a model output, and the copy-ready real INT4 deployment command. Interactive inference requires hosted GPU hardware; the model and loader remain fully available for local deployment.