Instructions to use TencentBAC/U-MARVEL-Qwen2VL-7B-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TencentBAC/U-MARVEL-Qwen2VL-7B-Instruct with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForSeq2SeqLM processor = AutoProcessor.from_pretrained("TencentBAC/U-MARVEL-Qwen2VL-7B-Instruct") model = AutoModelForSeq2SeqLM.from_pretrained("TencentBAC/U-MARVEL-Qwen2VL-7B-Instruct", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 3,819 Bytes
b22a67a 6b0b079 b22a67a 6b0b079 e91bc9d 6b0b079 7c31928 40784df 7c31928 40784df 7c31928 40784df 3d3d240 7c31928 3d3d240 7c31928 3d3d240 7c31928 4d99776 7c31928 4d99776 40784df 7c31928 40784df 7c31928 40784df 7c31928 40784df 7c31928 6b0b079 3d3d240 6b0b079 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 | ---
base_model:
- Qwen/Qwen2-VL-7B-Instruct
datasets:
- TIGER-Lab/M-BEIR
language:
- en
license: apache-2.0
pipeline_tag: any-to-any
library_name: transformers
---
## U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs
This repository contains the official model checkpoints and inference code for **U-MARVEL**, presented in the paper [U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs](https://huggingface.co/papers/2507.14902).
Universal multimodal retrieval (UMR) addresses complex retrieval tasks involving diverse modalities for both queries and candidates. Despite the success of state-of-the-art methods based on multimodal large language models (MLLMs) using contrastive learning principles, the mechanisms underlying their retrieval capabilities remain largely unexplored. This gap potentially leads to suboptimal performance and limited generalization ability.
In this study, we systematically analyze the key factors driving effective embedding learning for UMR using MLLMs. We implement a general MLLM-based embedding learning pipeline and investigate contributors to high-performing universal retrieval systems. Our analysis covers various aspects of embedding generation and training strategies, including progressive transition, hard negative mining, and re-ranker distillation. Our findings reveal that often-overlooked factors can significantly impact model performance.
Building on these insights, we introduce U-MARVEL (Universal Multimodal Retrieval via Embedding Learning), a unified framework that outperforms state-of-the-art competitors on the M-BEIR benchmark in supervised settings and demonstrates strong zero-shot performance on tasks such as composed image retrieval and text-to-video retrieval. These results highlight the generalization potential of our framework across various embedding-based retrieval tasks, providing valuable insights for future research.
## Model Checkpoints
```
βββ checkpoints
β βββ hf_models
β β βββ Qwen2-VL-7B-Instruct
β β βββ Qwen3-VL-4B-Instruct
β βββ U-MARVEL-Qwen2VL-7B-Instruct
β βββ U-MARVEL-Qwen3VL-4B-Instruct
```
- [U-MARVEL-Qwen2VL-7B-Instruct](https://huggingface.co/TencentBAC/U-MARVEL-Qwen2VL-7B-Instruct) π€
- [U-MARVEL-Qwen3VL-4B-Instruct](https://huggingface.co/TencentBAC/U-MARVEL-Qwen3VL-4B-Instruct) π€
- Code available at: [U-MARVEL](https://github.com/chaxjli/U-MARVEL)
## π Demo
To get started, first create a virtual environment and install the required dependencies:
```bash
conda create -n u-marvel python=3.9 -y
conda activate u-marvel
pip install -r requirements_qwen2_vl.txt
python demo.py
```
## Model Performance
The proposed U-MARVEL framework establishes new state-of-the-art performance across both
single-model architectures and recall-then-rerank approaches on M-BEIR benchmark.
<img src="./figures/local_pool.png" alt="M-BEIR-Local" width="700" height="auto">
<img src="./figures/global_pool.png" alt="M-BEIR-Global" width="700" height="auto">
<img src="./figures/zero_shot_image.png" alt="M-BEIR-Zero-shot_image" width="700" height="auto">
<img src="./figures/zero_shot_video.png" alt="M-BEIR-Zero-shot_video" width="700" height="auto">
## Acknowledgements
Many thanks to the code bases from **[LamRA](https://github.com/Code-kunkun/LamRA)** .
## Citation
If you use this code for your research or project, please cite:
```latex
@inproceedings{li2026umarvel,
title={U-{MARVEL}: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with {MLLM}s},
author={Xiaojie Li and Chu Li and Shi-Zhe Chen and Xi Chen},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026}
}
``` |