Spaces:
Sleeping
Sleeping
|
Download README.md from nanovdr/NanoVDR-Demo: direct link, hf CLI and curl.
- Browser
- Download file 2.5 kB
-
https://huggingface.co/spaces/nanovdr/NanoVDR-Demo/resolve/main/README.md
- Command line
-
hf download hf://spaces/nanovdr/NanoVDR-Demo/README.md
-
curl -L -o README.md https://huggingface.co/spaces/nanovdr/NanoVDR-Demo/resolve/main/README.md
2.5 kB
| title: NanoVDR Demo | |
| emoji: 🔍 | |
| colorFrom: blue | |
| colorTo: indigo | |
| sdk: gradio | |
| sdk_version: 6.28.0 | |
| app_file: app.py | |
| pinned: false | |
| license: apache-2.0 | |
| models: | |
| - nanovdr/NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML | |
| - nanovdr/ColNanoVDR-Q-Ettin150M-ColQwen35-320-ML | |
| - Qwen/Qwen3-VL-Embedding-2B | |
| - athrael-soju/colqwen3.5-4.5B-v3 | |
| datasets: | |
| - vidore/vidore_v3_computer_science | |
| tags: | |
| - visual-document-retrieval | |
| - late-interaction | |
| - nanovdr | |
| - colnanovdr | |
| # NanoVDR Interactive Retrieval Demo | |
| Asymmetric visual document retrieval: a small text-only encoder retrieves document pages indexed offline by a vision-language model, with no vision model at query time. Switch between the two students with the model selector. | |
| **Papers**: [NanoVDR](https://arxiv.org/abs/2603.12824) (single-vector) · [ColNanoVDR](https://arxiv.org/abs/2609.34899) (multi-vector) | |
| | | NanoVDR (single-vector) | ColNanoVDR (multi-vector) | | |
| |---|---|---| | |
| | Query encoder (CPU) | [NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML](https://huggingface.co/nanovdr/NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML), 69M | [ColNanoVDR-Q-Ettin150M-ColQwen35-320-ML](https://huggingface.co/nanovdr/ColNanoVDR-Q-Ettin150M-ColQwen35-320-ML), 149M | | |
| | Document index (offline) | page vectors of [Qwen3-VL-Embedding-2B](https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B) | page tokens of [ColQwen3.5-4.5B](https://huggingface.co/athrael-soju/colqwen3.5-4.5B-v3) | | |
| | Scoring | cosine similarity | MaxSim over weighted query tokens | | |
| **Corpus**: 1,360 pages from [ViDoRe v3 Computer Science](https://huggingface.co/datasets/vidore/vidore_v3_computer_science) (page images and queries under CC BY 4.0). For the example queries, each teacher's own top-k (pre-computed) and the annotated relevant pages are shown next to the student's results. | |
| ## Citation | |
| ```bibtex | |
| @article{nanovdr2026, | |
| title = {NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M | |
| Text-Only Encoder for Visual Document Retrieval}, | |
| author = {Liu, Zhuchenyang and Zhang, Yao and Xiao, Yu}, | |
| journal = {arXiv preprint arXiv:2603.12824}, | |
| year = {2026} | |
| } | |
| @misc{liu2026colnanovdrdocumentfreequerydistillation, | |
| title={ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport}, | |
| author={Zhuchenyang Liu and Ziyi Wang and Yao Zhang and Yu Xiao}, | |
| year={2026}, | |
| eprint={2609.34899}, | |
| archivePrefix={arXiv}, | |
| primaryClass={cs.IR}, | |
| url={https://arxiv.org/abs/2609.34899}, | |
| } | |
| ``` | |