Instructions to use nanovdr/ColNanoVDR-Q-Ettin150M-EVIE45B-2048-ML with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use nanovdr/ColNanoVDR-Q-Ettin150M-EVIE45B-2048-ML with sentence-transformers:
from sentence_transformers import MultiVectorEncoder model = MultiVectorEncoder("nanovdr/ColNanoVDR-Q-Ettin150M-EVIE45B-2048-ML") queries = ["Which planet is known as the Red Planet?"] documents = [ "Venus is often called Earth's twin because of its similar size and proximity.", "Mars, known for its reddish appearance, is often referred to as the Red Planet.", "Jupiter, the largest planet in our solar system, has a prominent red spot.", ] query_embeddings = model.encode_query(queries) document_embeddings = model.encode_document(documents) similarities = model.similarity(query_embeddings, document_embeddings) print(similarities) - Notebooks
- Google Colab
- Kaggle
ColNanoVDR: Document-Free Query-Side Distillation for Multi-Vector Visual Document Retrieval
Paper | Models | Code | Dataset
What is ColNanoVDR
ColNanoVDR replaces the multi-billion-parameter query encoder of a multi-vector visual document retriever with a small text-only encoder, and leaves the document side untouched. The student is trained by document-free distillation: it reproduces the teacher's query token embeddings under an optimal-transport objective (OTW), so no page image is ever encoded and no relevance label is used. At serving time the teacher's existing page index is scored by MaxSim unchanged, and with no vision tower on the query path a query can be encoded on CPU.
This tower targets tencent/EVIE-4.5B, a Prefix-MRL retriever whose single 2048-d projection can be truncated at serving time to 64, 128, 256, 512, 1024, or 2048 channels. The student is distilled with one OTW term per width, so one 150M tower serves every width the teacher serves, keeping 95.0% to 95.9% of the teacher's NDCG@5 on ViDoRe v3 at every width. Combined with EVIE's HAC index compression, the whole retrieval stack becomes a CPU query encoder and a few GiB of index per million pages.
This model
Every ColNanoVDR release is named:
ColNanoVDR-Q-{backbone}-{teacher}-{dim}-ML
| Field | Meaning | This model |
|---|---|---|
Q |
the query tower: the only side that is trained; documents stay with the teacher | Q |
{backbone} |
the text-only encoder distilled into | Ettin150M = jhu-clsp/ettin-encoder-150m |
{teacher} |
the multi-vector VDR teacher whose embedding space this tower targets | EVIE45B = tencent/EVIE-4.5B, 4.5B params |
{dim} |
late-interaction vector width of that teacher | 2048 (Prefix-MRL: also 64, 128, 256, 512, 1024) |
ML |
multilingual training mixture | English + 5 Latin-script European languages |
A tower is only valid with the teacher it was distilled from, and query and pages must use the same width. This one requires pages indexed by tencent/EVIE-4.5B. Towers for other teachers are listed on the nanovdr page.
Training
Dataset. nanovdr/NanoVDR-Train: 1.49M queries, 711K English plus 778K machine translations into five Latin-script European languages. Only the query text is used. EVIE-4.5B query embeddings are cached once at full width, so the teacher never runs during training.
Loss: OTW per prefix. A query is a set of token vectors, and the student and the
teacher tokenize differently, so there is no token-to-token correspondence. OTW treats
each side as a weighted set on the unit sphere and minimises the entropic
optimal-transport cost between them under c(s,t) = 1 - <s,t>, with a learned weight per
student token. For a Prefix-MRL teacher, both sides are truncated to each width d and
renormalised, exactly as the teacher does at serving time, and one OTW cost is computed
per width with the same learned weights:
L = sum_d w_d <P_d, C_d>, C_d = 1 - <s_i[:d]/|s_i[:d]|, t_j[:d]/|t_j[:d]|>,
d in {64, 128, 256, 512, 1024, 2048}, w_d uniform
Settings: eps = 0.05, 50 Sinkhorn iterations, AdamW one-cycle, peak LR 3e-4, 3% warmup,
effective batch 1024 (128 per GPU x 2 GPUs x 4 accumulation), 10 epochs.
Architecture. Ettin encoder -> bias-free linear projection to 2048-d -> per-token L2
normalisation -> learned weight head (a linear layer over the pre-projection hidden
states, softmax over the query's tokens). At width d each token is truncated to d
channels, renormalised, and scaled by its weight, so the emitted vectors are deliberately
not unit-length and their norms sum to 1 per query. Special tokens are excluded from
scoring. 150M parameters.
Usage
Requires sentence-transformers>=6.0 and transformers>=5.0 (tested with 6.1 and
5.13.1). Retrieval runs in two stages: the teacher indexes your pages once, offline;
ColNanoVDR answers every query online. Pick one width D and use it on both sides.
Step 1 (offline, once): index your pages with EVIE-4.5B at width D
EVIE-4.5B ships its own Sentence Transformers integration. Encode pages at full width and
keep the renormalised D-channel prefix, as EVIE does at serving time:
import torch
import torch.nn.functional as F
from sentence_transformers import MultiVectorEncoder
D = 128
teacher = MultiVectorEncoder("tencent/EVIE-4.5B", device="cuda",
model_kwargs={"torch_dtype": torch.bfloat16})
pages = teacher.encode_document(page_images, convert_to_numpy=False) # (n_tokens, 2048) each
page_embeddings = [F.normalize(p.float()[:, :D], dim=-1) for p in pages]
This gives the same page embeddings as EVIE's own colpali_engine code with
set_active_head(D) and 1024 visual tokens per page, the setting the numbers below were
measured at. Optionally compress the index with EVIE's HAC (code/compress in
Tencent/EVIE).
Step 2 (online, per query): encode with ColNanoVDR at the same width
from sentence_transformers import MultiVectorEncoder
model = MultiVectorEncoder("nanovdr/ColNanoVDR-Q-Ettin150M-EVIE45B-2048-ML", trust_remote_code=True)
model[len(model) - 1].set_prefix_dim(D) # serve width D (None = full 2048)
q = model.encode_query(["What was the revenue growth in Q3 2024?"])
scores = model.similarity(q, page_embeddings) # meanMaxSim
No teacher, no image processor, no GPU required on this path. Notes that change results if ignored:
- Set the width with
set_prefix_dim, never by slicing the output. The learned weights are folded into the vector norms; truncating and renormalising the emitted vectors would erase them.set_prefix_dimrenormalises the prefix first and applies the weights afterwards. - Pass the raw query, with no instruction prefix, and do not re-normalise the output.
Performance
ViDoRe v3, the 8 public datasets, all query languages. Pages are encoded by EVIE-4.5B at 1024 visual tokens; the teacher row scores them with EVIE's own queries, the student row with ColNanoVDR's, so all error is attributed to the query tower. Measured with our harness; EVIE-4.5B at 2048 scores 65.86 NDCG@10 here against 66.02 in EVIE's own report.
| Width | Bytes / vector | EVIE-4.5B NDCG@5 | ColNanoVDR NDCG@5 | Retention | EVIE-4.5B NDCG@10 | ColNanoVDR NDCG@10 | Retention |
|---|---|---|---|---|---|---|---|
| 64 | 128 B | 61.86 | 58.76 | 95.0% | 64.28 | 61.18 | 95.2% |
| 128 | 256 B | 62.54 | 59.83 | 95.7% | 65.07 | 62.29 | 95.7% |
| 256 | 512 B | 62.89 | 60.32 | 95.9% | 65.45 | 62.83 | 96.0% |
| 512 | 1,024 B | 63.17 | 60.40 | 95.6% | 65.69 | 62.87 | 95.7% |
| 1024 | 2,048 B | 63.20 | 60.32 | 95.4% | 65.76 | 62.71 | 95.4% |
| 2048 | 4,096 B | 63.25 | 60.19 | 95.2% | 65.86 | 62.68 | 95.2% |
With HAC index compression
EVIE ships a training-free Hierarchical Agglomerative Clustering (HAC) step that pools
each page's image tokens into a fixed number of centroids at a chosen width. Applying
it with EVIE's own code (code/compress, position weight 0.1) and scoring the same
compressed index with the teacher's queries and with ColNanoVDR:
| Index | Vectors / page | Index size / 1M pages | EVIE-4.5B NDCG@10 | ColNanoVDR NDCG@10 | Retention |
|---|---|---|---|---|---|
| d64 K32 | 32.0 | 3.81 GiB | 59.44 | 55.76 | 93.8% |
| d64 K64 | 64.0 | 7.63 GiB | 61.82 | 58.49 | 94.6% |
| d128 K32 | 32.0 | 7.63 GiB | 61.21 | 58.07 | 94.9% |
EVIE's own report gives 59.58, 62.06 and 61.40 NDCG@10 for the teacher on these three indexes; our runs of its code reproduce them to within 0.24 points. The query tower needs no change for a compressed index: it is the same tower at the same width.
Efficiency
| Params | Vision tower at query time | CPU serving | |
|---|---|---|---|
| EVIE-4.5B (teacher) | 4.5B | required | no |
| ColNanoVDR-Q-Ettin150M-EVIE45B-2048-ML | 150M | none | yes |
License
Apache-2.0, matching the teacher, tencent/EVIE-4.5B (Apache-2.0). The query tower was distilled from EVIE-4.5B's query embeddings; please also credit EVIE (citation below). The backbone, Ettin, is MIT.
Contact
Zhuchenyang Liu, Aalto University: zhuchenyang.liu@aalto.fi
Citation
If you use ColNanoVDR, please cite:
@misc{liu2026colnanovdrdocumentfreequerydistillation,
title={ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport},
author={Zhuchenyang Liu and Ziyi Wang and Yao Zhang and Yu Xiao},
year={2026},
eprint={2609.34899},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2609.34899},
}
The teacher:
@misc{tencent2026evie,
title = {EVIE: High-Performance Multilingual Visual Document Retrieval with Matryoshka Embeddings and Token Compression},
author = {Wang, Zifei and Wen, Wei},
year = {2026},
howpublished = {\url{https://github.com/Tencent/EVIE}},
note = {Corresponding author: Wei Wen <jawnrwen@tencent.com>}
}
- Downloads last month
- 25
Model tree for nanovdr/ColNanoVDR-Q-Ettin150M-EVIE45B-2048-ML
Base model
jhu-clsp/ettin-encoder-150m