Instructions to use nanovdr/NanoVDR-D-HiRes-Qwen3VL8B-4096 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use nanovdr/NanoVDR-D-HiRes-Qwen3VL8B-4096 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("nanovdr/NanoVDR-D-HiRes-Qwen3VL8B-4096", trust_remote_code=True) sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Transformers
How to use nanovdr/NanoVDR-D-HiRes-Qwen3VL8B-4096 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="nanovdr/NanoVDR-D-HiRes-Qwen3VL8B-4096", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("nanovdr/NanoVDR-D-HiRes-Qwen3VL8B-4096", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
NanoVDR-D-HiRes-Qwen3VL8B-4096: the document tower
Models | Code | Demo | Write-up | Paper
Renamed. This model was published as
nanovdr/NanoVDR-D-HiRes. Old links redirect. The name now states the tower (D), the tile budget, the teacher it was distilled from and the output width: pair it with any tower whose teacher and width match, in this case NanoVDR-Q-DistilBERT-Qwen3VL8B-4096-ML.
A 457M encoder that turns a document page image into one 4096-d vector. It is distilled from a frozen 8B vision-language teacher (Qwen3-VL-Embedding-8B) by direct representation alignment: the student is trained only to reproduce the teacher's page embedding, with no relevance labels, no negative mining and no contrastive term.
Every other NanoVDR release replaces the query side and still needs the teacher to build the index. This one replaces the document side, so paired with a NanoVDR query tower the teacher is not needed at all, at indexing time or at query time.
| Input | page image (PIL / any resolution) |
| Output | one L2-normalised 4096-d vector |
| Scoring | dot product |
| Index footprint | 16.4 GB per million pages (float32) |
| Teacher | Qwen3-VL-Embedding-8B, frozen |
| Visual tokens | up to 6 tiles at 448x448 + one whole-page thumbnail |
How it works
page ──► dynamic tiling ──► InternViT-300M-448 ──► linear ──► ModernBERT-base ──► mean pool ──► Linear(768→4096) ──► L2-norm
(≤6 tiles + per-tile patch 1024→768 bidirectional,
thumbnail) tokens, concatenated no text tokens
The tiling rule matches the page's aspect ratio to a grid at the encoder's native 448px resolution and appends a thumbnail for global context, so small text and table cells survive without the patch sequence outrunning ModernBERT's 8192-token window.
Usage
Both towers are called the same way. The document tower needs
sentence-transformers>=5.4, which is the release that routes image inputs to
the model.
from PIL import Image
from sentence_transformers import SentenceTransformer
docs = SentenceTransformer("nanovdr/NanoVDR-D-HiRes-Qwen3VL8B-4096", trust_remote_code=True)
queries = SentenceTransformer("nanovdr/NanoVDR-Q-DistilBERT-Qwen3VL8B-4096-ML")
doc_emb = docs.encode([Image.open("page1.png"), Image.open("page2.png")])
query_emb = queries.encode(["what was the revenue growth in Q3 2024?"])
scores = query_emb @ doc_emb.T # (n_queries, n_pages)
Through transformers instead, which has no version floor:
import torch
from transformers import AutoModel, AutoImageProcessor
model = AutoModel.from_pretrained("nanovdr/NanoVDR-D-HiRes-Qwen3VL8B-4096", trust_remote_code=True).eval()
processor = AutoImageProcessor.from_pretrained("nanovdr/NanoVDR-D-HiRes-Qwen3VL8B-4096", trust_remote_code=True)
inputs = processor(images=pages, return_tensors="pt") # pixel_values + tile_mask
with torch.no_grad():
doc_emb = model(**inputs).embedding # (N, 4096), L2-normalised
processor is a NanoVDRDocImageProcessor. It tiles the page and returns the
tile_mask the model needs next to pixel_values, so a stock single-view image
processor is not a substitute: it silently costs several NDCG points, and the
model warns if it is handed one. model.encode(pages, processor, batch_size=4)
does the same thing with batching.
On Ampere or newer, switch the fused kernels on:
model = AutoModel.from_pretrained(
"nanovdr/NanoVDR-D-HiRes-Qwen3VL8B-4096",
trust_remote_code=True,
use_flash_attn=True, # InternViT
text_attn_implementation="flash_attention_2",
).to("cuda", torch.bfloat16)
Weights ship in float32 so the model loads on any device; cast to bfloat16 for deployment.
Any half of the teacher, swapped
Both towers are trained to land in the same frozen teacher's embedding space, so either one is a drop-in replacement for the corresponding half of that teacher. All four combinations are valid retrieval systems, and they trade quality against what you no longer have to run:
| teacher documents | student documents | |
|---|---|---|
| teacher queries | 71.05. The ceiling, 8B on both sides | 65.02. Indexing is 7x cheaper, teacher still needed per query |
| student queries | 66.36. Queries encode on one CPU thread, index built once by the teacher | 61.74. No teacher anywhere |
Average NDCG@5 over ViDoRe v1+v2+v3, with NanoVDR-D-HiRes-Qwen3VL8B-4096 and NanoVDR-Q-DistilBERT-Qwen3VL8B-4096-ML against Qwen3-VL-Embedding-8B.
The pairing rule follows from the naming: the teacher and the dimension must match, and any cell of the matrix is then available. The teacher fixes the space; the dimension fixes which width of it is targeted, since a Matryoshka teacher can be aligned to at several widths.
Results
NDCG@5 on all 22 datasets of ViDoRe v1, v2 and v3. Every baseline was reproduced locally under one protocol rather than quoted from its own paper, so these numbers are comparable to each other and not to published tables.
End to end, paired with the 70M text-only query tower distilled from the same teacher, 527M total, with nothing multi-billion running at deployment:
| Model | Params | Type | v1 | v2 | v3 | Avg |
|---|---|---|---|---|---|---|
| NanoVDR-D-HiRes-Qwen3VL8B-4096 + 70M query tower | 527M | single | 82.81 | 55.34 | 47.07 | 61.74 |
| colSmol-500M | 478M | multi | 82.42 | 43.09 | 33.52 | 53.01 |
| SauerkrautLM-ColLFM2 | 451M | multi | 78.24 | 45.09 | 33.19 | 52.17 |
| colSmol-256M | 256M | multi | 79.72 | 34.63 | 25.23 | 46.53 |
| ColModernVBERT | 250M | multi | 76.76 | 33.18 | 17.45 | 42.46 |
| SigLIP2-L | 880M | single | 43.58 | 20.17 | 14.04 | 25.93 |
| BiModernVBERT | 250M | single | 37.40 | 10.88 | 5.52 | 17.93 |
| Qwen3-VL-Embedding-8B (teacher) | 8.1B | single | 87.31 | 69.76 | 56.07 | 71.05 |
That is 86.9% of the teacher, and a lead of 8.73 average points over the strongest sub-1B baseline we reproduced.
Document side in isolation, scored against teacher-encoded queries, which removes query-side error:
| v1 | v2 | v3 | Avg | |
|---|---|---|---|---|
| NanoVDR-D-HiRes-Qwen3VL8B-4096 | 83.72 | 60.92 | 50.43 | 65.02 |
Deployment cost
Single H200, batch size 8, bfloat16. "Score 10K" is one query against 10 000 candidates on a single CPU thread.
| this model | colSmol-500M | |
|---|---|---|
| Document throughput | 36.82 pages/s | 2.82 pages/s |
| Peak VRAM | 3.07 GB | 4.97 GB |
| Index per 1M pages | 16.4 GB | 256 GB |
| Score 10K | 9.6 ms | 1 161 ms |
One vector per page instead of a token set is what buys the last two rows: a 15.6x smaller index and two orders of magnitude less scoring work than the multi-vector retrievers in the same parameter class.
Training
| Objective | 1 - cos(student, teacher), nothing else |
| Data | 1.20M unique page images, public and permissively licensed |
| Dedup | perceptual hash against all three ViDoRe evaluation corpora |
| Optimizer | AdamW, one-cycle, peak LR 1e-3, 3% warmup |
| Effective batch | 256 |
| Epochs | 3 |
| Hardware | 2x H200 |
Teacher targets are cached once before training, so the document and query towers train independently and in parallel.
Limitations
- Bounded by the teacher. The student only reproduces Qwen3-VL-Embedding-8B's embedding space; any systematic weakness of that teacher carries over, and nothing in the objective lets the student exceed it.
- Fixed tile budget. Tiling adapts to page aspect ratio but not to page content, so a dense small-text page and a single-figure page get the same visual-token budget.
- Evaluation is ViDoRe-only. Three difficulty levels, six languages, eight professional domains, but not in-house enterprise layouts, scanned or OCR-degraded pages, or production query distributions.
- Uncompressed index. One 4096-d float32 vector per page. Quantisation and product quantisation are untested here, so 16.4 GB per million is an upper bound.
Citation
This model is described in DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation.
@article{distilvdr2026,
title = {DistilVDR: A Compact End-to-End Visual Document Retriever
via Dual-Student Distillation},
author = {Liu, Zhuchenyang and Wang, Ziyi and Zhang, Yao and Xiao, Yu},
journal = {arXiv preprint arXiv:2608.10636},
year = {2026}
}
Distilling the Document Tower is the write-up behind this model: the teacher and data choices, the architecture, all fourteen ablations, and the five things that did not work.
License
MIT. The visual and text backbones follow the licences of InternViT-300M-448px-V2_5 and ModernBERT-base.
- Downloads last month
- 40
Model tree for nanovdr/NanoVDR-D-HiRes-Qwen3VL8B-4096
Base model
OpenGVLab/InternViT-300M-448px