NanoVDR-Demo / README.md
Ryenhails's picture
List NanoVDR first; NanoVDR is the default model
1184bf3 verified
|
Raw History Blame Contribute Delete
2.5 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade
metadata
title: NanoVDR Demo
emoji: 🔍
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 6.28.0
app_file: app.py
pinned: false
license: apache-2.0
models:
  - nanovdr/NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML
  - nanovdr/ColNanoVDR-Q-Ettin150M-ColQwen35-320-ML
  - Qwen/Qwen3-VL-Embedding-2B
  - athrael-soju/colqwen3.5-4.5B-v3
datasets:
  - vidore/vidore_v3_computer_science
tags:
  - visual-document-retrieval
  - late-interaction
  - nanovdr
  - colnanovdr

NanoVDR Interactive Retrieval Demo

Asymmetric visual document retrieval: a small text-only encoder retrieves document pages indexed offline by a vision-language model, with no vision model at query time. Switch between the two students with the model selector.

Papers: NanoVDR (single-vector) · ColNanoVDR (multi-vector)

NanoVDR (single-vector) ColNanoVDR (multi-vector)
Query encoder (CPU) NanoVDR-Q-DistilBERT-Qwen3VL2B-2048-ML, 69M ColNanoVDR-Q-Ettin150M-ColQwen35-320-ML, 149M
Document index (offline) page vectors of Qwen3-VL-Embedding-2B page tokens of ColQwen3.5-4.5B
Scoring cosine similarity MaxSim over weighted query tokens

Corpus: 1,360 pages from ViDoRe v3 Computer Science (page images and queries under CC BY 4.0). For the example queries, each teacher's own top-k (pre-computed) and the annotated relevant pages are shown next to the student's results.

Citation

@article{nanovdr2026,
  title   = {NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M
             Text-Only Encoder for Visual Document Retrieval},
  author  = {Liu, Zhuchenyang and Zhang, Yao and Xiao, Yu},
  journal = {arXiv preprint arXiv:2603.12824},
  year    = {2026}
}

@misc{liu2026colnanovdrdocumentfreequerydistillation,
      title={ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport}, 
      author={Zhuchenyang Liu and Ziyi Wang and Yao Zhang and Yu Xiao},
      year={2026},
      eprint={2609.34899},
      archivePrefix={arXiv},
      primaryClass={cs.IR},
      url={https://arxiv.org/abs/2609.34899}, 
}