metadata
license: apache-2.0
base_model:
- Qwen/Qwen2.5-VL-3B-Instruct
tags:
- vision-language-model
- efficient-vlm
- llava
- qwen2-vl
- token-compression
- frequency-domain
Fourier-Qwen2.5-VL-3B-0.67
Official checkpoints for Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models.
Model Details
| Model | Base Model | Visual Tokens | Compression | Weights |
|---|---|---|---|---|
| Fourier-LLaVA-v1.5-7B-256 | LLaVA-v1.5-7B | 256 | 55.6% | 🤗 HF |
| Fourier-LLaVA-v1.5-7B-144 | LLaVA-v1.5-7B | 144 | 75.0% | 🤗 HF |
| Fourier-LLaVA-v1.5-7B-64 | LLaVA-v1.5-7B | 64 | 88.9% | 🤗 HF |
| Fourier-LLaVA-v1.5-7B-36 | LLaVA-v1.5-7B | 36 | 93.8% | 🤗 HF |
| Fourier-LLaVA-v1.5-13B-144 | LLaVA-v1.5-13B | 144 | 75.0% | 🤗 HF |
| Fourier-Qwen2-VL-2B-0.67 | Qwen2-VL-2B-Instruct | Dynamic | 55.6% | 🤗 HF |
| Fourier-Qwen2.5-VL-3B-0.67 | Qwen2.5-VL-3B-Instruct | Dynamic | 55.6% | 🤗 HF |