ColNanoVDR: Document-Free Query-Side Distillation for Multi-Vector Visual Document Retrieval

Paper  |  Models  |  Code  |  Dataset


What is ColNanoVDR

ColNanoVDR replaces the multi-billion-parameter query encoder of a multi-vector visual document retriever with a small text-only encoder, and leaves the document side untouched. The student is trained by document-free distillation: it reproduces the teacher's query token embeddings under an optimal-transport objective (OTW), so no page image is ever encoded and no relevance label is used. At serving time the teacher's existing page index is scored by MaxSim unchanged, and with no vision tower on the query path a query can be encoded on CPU.

This tower targets tencent/EVIE-4.5B, a Prefix-MRL retriever whose single 2048-d projection can be truncated at serving time to 64, 128, 256, 512, 1024, or 2048 channels. The student is distilled with one OTW term per width, so one 150M tower serves every width the teacher serves, keeping 95.0% to 95.9% of the teacher's NDCG@5 on ViDoRe v3 at every width. Combined with EVIE's HAC index compression, the whole retrieval stack becomes a CPU query encoder and a few GiB of index per million pages.

This model

Every ColNanoVDR release is named:

ColNanoVDR-Q-{backbone}-{teacher}-{dim}-ML
Field Meaning This model
Q the query tower: the only side that is trained; documents stay with the teacher Q
{backbone} the text-only encoder distilled into Ettin150M = jhu-clsp/ettin-encoder-150m
{teacher} the multi-vector VDR teacher whose embedding space this tower targets EVIE45B = tencent/EVIE-4.5B, 4.5B params
{dim} late-interaction vector width of that teacher 2048 (Prefix-MRL: also 64, 128, 256, 512, 1024)
ML multilingual training mixture English + 5 Latin-script European languages

A tower is only valid with the teacher it was distilled from, and query and pages must use the same width. This one requires pages indexed by tencent/EVIE-4.5B. Towers for other teachers are listed on the nanovdr page.

Training

Dataset. nanovdr/NanoVDR-Train: 1.49M queries, 711K English plus 778K machine translations into five Latin-script European languages. Only the query text is used. EVIE-4.5B query embeddings are cached once at full width, so the teacher never runs during training.

Loss: OTW per prefix. A query is a set of token vectors, and the student and the teacher tokenize differently, so there is no token-to-token correspondence. OTW treats each side as a weighted set on the unit sphere and minimises the entropic optimal-transport cost between them under c(s,t) = 1 - <s,t>, with a learned weight per student token. For a Prefix-MRL teacher, both sides are truncated to each width d and renormalised, exactly as the teacher does at serving time, and one OTW cost is computed per width with the same learned weights:

L = sum_d w_d <P_d, C_d>,   C_d = 1 - <s_i[:d]/|s_i[:d]|, t_j[:d]/|t_j[:d]|>,
d in {64, 128, 256, 512, 1024, 2048},  w_d uniform

Settings: eps = 0.05, 50 Sinkhorn iterations, AdamW one-cycle, peak LR 3e-4, 3% warmup, effective batch 1024 (128 per GPU x 2 GPUs x 4 accumulation), 10 epochs.

Architecture. Ettin encoder -> bias-free linear projection to 2048-d -> per-token L2 normalisation -> learned weight head (a linear layer over the pre-projection hidden states, softmax over the query's tokens). At width d each token is truncated to d channels, renormalised, and scaled by its weight, so the emitted vectors are deliberately not unit-length and their norms sum to 1 per query. Special tokens are excluded from scoring. 150M parameters.

Usage

Requires sentence-transformers>=6.0 and transformers>=5.0 (tested with 6.1 and 5.13.1). Retrieval runs in two stages: the teacher indexes your pages once, offline; ColNanoVDR answers every query online. Pick one width D and use it on both sides.

Step 1 (offline, once): index your pages with EVIE-4.5B at width D

EVIE-4.5B ships its own Sentence Transformers integration. Encode pages at full width and keep the renormalised D-channel prefix, as EVIE does at serving time:

import torch
import torch.nn.functional as F
from sentence_transformers import MultiVectorEncoder

D = 128
teacher = MultiVectorEncoder("tencent/EVIE-4.5B", device="cuda",
                             model_kwargs={"torch_dtype": torch.bfloat16})
pages = teacher.encode_document(page_images, convert_to_numpy=False)   # (n_tokens, 2048) each
page_embeddings = [F.normalize(p.float()[:, :D], dim=-1) for p in pages]

This gives the same page embeddings as EVIE's own colpali_engine code with set_active_head(D) and 1024 visual tokens per page, the setting the numbers below were measured at. Optionally compress the index with EVIE's HAC (code/compress in Tencent/EVIE).

Step 2 (online, per query): encode with ColNanoVDR at the same width

from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder("nanovdr/ColNanoVDR-Q-Ettin150M-EVIE45B-2048-ML", trust_remote_code=True)
model[len(model) - 1].set_prefix_dim(D)         # serve width D (None = full 2048)

q = model.encode_query(["What was the revenue growth in Q3 2024?"])
scores = model.similarity(q, page_embeddings)   # meanMaxSim

No teacher, no image processor, no GPU required on this path. Notes that change results if ignored:

  • Set the width with set_prefix_dim, never by slicing the output. The learned weights are folded into the vector norms; truncating and renormalising the emitted vectors would erase them. set_prefix_dim renormalises the prefix first and applies the weights afterwards.
  • Pass the raw query, with no instruction prefix, and do not re-normalise the output.

Performance

ViDoRe v3, the 8 public datasets, all query languages. Pages are encoded by EVIE-4.5B at 1024 visual tokens; the teacher row scores them with EVIE's own queries, the student row with ColNanoVDR's, so all error is attributed to the query tower. Measured with our harness; EVIE-4.5B at 2048 scores 65.86 NDCG@10 here against 66.02 in EVIE's own report.

Width Bytes / vector EVIE-4.5B NDCG@5 ColNanoVDR NDCG@5 Retention EVIE-4.5B NDCG@10 ColNanoVDR NDCG@10 Retention
64 128 B 61.86 58.76 95.0% 64.28 61.18 95.2%
128 256 B 62.54 59.83 95.7% 65.07 62.29 95.7%
256 512 B 62.89 60.32 95.9% 65.45 62.83 96.0%
512 1,024 B 63.17 60.40 95.6% 65.69 62.87 95.7%
1024 2,048 B 63.20 60.32 95.4% 65.76 62.71 95.4%
2048 4,096 B 63.25 60.19 95.2% 65.86 62.68 95.2%

With HAC index compression

EVIE ships a training-free Hierarchical Agglomerative Clustering (HAC) step that pools each page's image tokens into a fixed number of centroids at a chosen width. Applying it with EVIE's own code (code/compress, position weight 0.1) and scoring the same compressed index with the teacher's queries and with ColNanoVDR:

Index Vectors / page Index size / 1M pages EVIE-4.5B NDCG@10 ColNanoVDR NDCG@10 Retention
d64 K32 32.0 3.81 GiB 59.44 55.76 93.8%
d64 K64 64.0 7.63 GiB 61.82 58.49 94.6%
d128 K32 32.0 7.63 GiB 61.21 58.07 94.9%

EVIE's own report gives 59.58, 62.06 and 61.40 NDCG@10 for the teacher on these three indexes; our runs of its code reproduce them to within 0.24 points. The query tower needs no change for a compressed index: it is the same tower at the same width.

Efficiency

Params Vision tower at query time CPU serving
EVIE-4.5B (teacher) 4.5B required no
ColNanoVDR-Q-Ettin150M-EVIE45B-2048-ML 150M none yes

License

Apache-2.0, matching the teacher, tencent/EVIE-4.5B (Apache-2.0). The query tower was distilled from EVIE-4.5B's query embeddings; please also credit EVIE (citation below). The backbone, Ettin, is MIT.

Contact

Zhuchenyang Liu, Aalto University: zhuchenyang.liu@aalto.fi

Citation

If you use ColNanoVDR, please cite:

@misc{liu2026colnanovdrdocumentfreequerydistillation,
      title={ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport}, 
      author={Zhuchenyang Liu and Ziyi Wang and Yao Zhang and Yu Xiao},
      year={2026},
      eprint={2609.34899},
      archivePrefix={arXiv},
      primaryClass={cs.IR},
      url={https://arxiv.org/abs/2609.34899}, 
}

The teacher:

@misc{tencent2026evie,
  title        = {EVIE: High-Performance Multilingual Visual Document Retrieval with Matryoshka Embeddings and Token Compression},
  author       = {Wang, Zifei and Wen, Wei},
  year         = {2026},
  howpublished = {\url{https://github.com/Tencent/EVIE}},
  note         = {Corresponding author: Wei Wen <jawnrwen@tencent.com>}
}
Downloads last month
25
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nanovdr/ColNanoVDR-Q-Ettin150M-EVIE45B-2048-ML

Finetuned
(29)
this model

Dataset used to train nanovdr/ColNanoVDR-Q-Ettin150M-EVIE45B-2048-ML

Paper for nanovdr/ColNanoVDR-Q-Ettin150M-EVIE45B-2048-ML