NanoVDR-Demo / README.md
Ryenhails's picture
List NanoVDR first; NanoVDR is the default model
1184bf3 verified
|
Raw History Blame Contribute Delete
2.5 kB
---
title: NanoVDR Demo
emoji: 🔍
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 6.28.0
app_file: app.py
pinned: false
license: apache-2.0
models:
- nanovdr/NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML
- nanovdr/ColNanoVDR-Q-Ettin150M-ColQwen35-320-ML
- Qwen/Qwen3-VL-Embedding-2B
- athrael-soju/colqwen3.5-4.5B-v3
datasets:
- vidore/vidore_v3_computer_science
tags:
- visual-document-retrieval
- late-interaction
- nanovdr
- colnanovdr
---
# NanoVDR Interactive Retrieval Demo
Asymmetric visual document retrieval: a small text-only encoder retrieves document pages indexed offline by a vision-language model, with no vision model at query time. Switch between the two students with the model selector.
**Papers**: [NanoVDR](https://arxiv.org/abs/2603.12824) (single-vector) · [ColNanoVDR](https://arxiv.org/abs/2609.34899) (multi-vector)
| | NanoVDR (single-vector) | ColNanoVDR (multi-vector) |
|---|---|---|
| Query encoder (CPU) | [NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML](https://huggingface.co/nanovdr/NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML), 69M | [ColNanoVDR-Q-Ettin150M-ColQwen35-320-ML](https://huggingface.co/nanovdr/ColNanoVDR-Q-Ettin150M-ColQwen35-320-ML), 149M |
| Document index (offline) | page vectors of [Qwen3-VL-Embedding-2B](https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B) | page tokens of [ColQwen3.5-4.5B](https://huggingface.co/athrael-soju/colqwen3.5-4.5B-v3) |
| Scoring | cosine similarity | MaxSim over weighted query tokens |
**Corpus**: 1,360 pages from [ViDoRe v3 Computer Science](https://huggingface.co/datasets/vidore/vidore_v3_computer_science) (page images and queries under CC BY 4.0). For the example queries, each teacher's own top-k (pre-computed) and the annotated relevant pages are shown next to the student's results.
## Citation
```bibtex
@article{nanovdr2026,
title = {NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M
Text-Only Encoder for Visual Document Retrieval},
author = {Liu, Zhuchenyang and Zhang, Yao and Xiao, Yu},
journal = {arXiv preprint arXiv:2603.12824},
year = {2026}
}
@misc{liu2026colnanovdrdocumentfreequerydistillation,
title={ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport},
author={Zhuchenyang Liu and Ziyi Wang and Yao Zhang and Yu Xiao},
year={2026},
eprint={2609.34899},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2609.34899},
}
```