--- library_name: sentence-transformers license: apache-2.0 language: - multilingual base_model: - Hcompany/NeoMME-800M-Retriever pipeline_tag: sentence-similarity tags: - multimodal - document-retrieval - dense-retrieval ---

NeoMME logo

# NeoMME-Retriever (800M): Single-Tower Multimodal-Native Multilingual Foundation Encoder 🔎 > [!IMPORTANT] > NeoMME-Retriever (800M) variants: > > - [Default (`transformers`)](https://huggingface.co/Hcompany/NeoMME-800M-Retriever): Returns dense and multi-vector embeddings together with a single forward pass. Recommended for most use cases and inference. > - [ST dense](https://huggingface.co/Hcompany/NeoMME-800M-Retriever-ST-dense) [current]: Supports independent dense fine-tuning with Sentence Transformers. > - [ST late-interaction](https://huggingface.co/Hcompany/NeoMME-800M-Retriever-ST-late): Supports independent multi-vector fine-tuning with Sentence Transformers. [![Hugging Face](https://img.shields.io/badge/Model_doc-FFD21E?style=for-the-badge&logo=huggingface&logoColor=000)](https://huggingface.co/docs/transformers/en/model_doc/neomme) [![Hugging Face](https://img.shields.io/badge/Collection-FFD21E?style=for-the-badge&logo=huggingface&logoColor=000)](https://hf.co/collections/Hcompany/neomme) [![arXiv](https://img.shields.io/badge/arXiv-2609.01657-b31b1b.svg?style=for-the-badge)](https://arxiv.org/abs/2609.01657) NeoMME-800M-Retriever-ST-dense is a model for multimodal document retrieval. Fine-tuned from [NeoMME-800M](https://huggingface.co/Hcompany/NeoMME-800M), it encodes text queries and documents (text or page screenshots) using one shared bidirectional Transformer encoder. This model can be used with Sentence Transformers, but can only generate dense embeddings.
SpecificationValue
Parameters800M
Vocabulary131,072 tokens
Context length16,384 tokens
Hidden size1,792
Image patches32 × 32 pixels, up to 2,048 pixels on the longest side (default)
Dense embeddings1,792 dimensions (Matryoshka: [128, 256, 512, 1,024, 1,792])
Dense pooling strategyMean
Dense embeddings are L2-normalized and use cosine similarity. They match `NeoMMEForRetrieval.dense_embeddings`. ## Performance All scores use the metric shown at the full trained dimensions. Higher is better. ViDoRe v3, v2, and v1 measure visual document retrieval, while BEIR-15 measures text retrieval.
BenchmarkMetricNeoMME-260MNeoMME-800M
Late interactionDenseLate interactionDense [current]
ViDoRe v3nDCG@100.52260.39070.55600.4391
ViDoRe v2nDCG@50.52180.40750.55910.4475
ViDoRe v1nDCG@50.85980.75520.87440.7993
BEIR-15nDCG@100.48810.30550.51260.3686
## Usage ```bash pip install -U "sentence-transformers[image]" ``` ```python from sentence_transformers import SentenceTransformer model = SentenceTransformer("Hcompany/NeoMME-800M-Retriever-ST-dense") queries = [ "Quelle partie de la production pétrolière du Kazakhstan provient de champs en mer ?", "Which hour of the day had the highest overall electricity generation in 2019?", ] documents = [ "https://github.com/tonywu71/colpali-cookbooks/blob/main/examples/data/shift_kazakhstan.jpg?raw=true", "https://github.com/tonywu71/colpali-cookbooks/blob/main/examples/data/energy_electricity_generation.jpg?raw=true", ] query_embeddings = model.encode_query(queries, convert_to_tensor=True) document_embeddings = model.encode_document(documents, convert_to_tensor=True) scores = model.similarity(query_embeddings, document_embeddings) # Expected: scores[0, 0] > scores[0, 1] and scores[1, 1] > scores[1, 0]. print(scores) ``` The score tensor has shape `(num_queries, num_documents)` and `scores[i, j]` is the score between query `i` and document `j`. A larger value indicates a closer match. ## Training NeoMME-800M-Retriever was fine-tuned from [NeoMME-800M](https://huggingface.co/Hcompany/NeoMME-800M) on text retrieval and document-page images. Training uses a joint late-interaction and Matryoshka dense contrastive objective. The [NeoMME technical report](https://arxiv.org/abs/2609.01657) describes the full fine-tuning recipe. ## Limitations With Sentence Transformers, only one of the two retrieval heads can be used at a time. ## License Model weights are released under the Apache 2.0 license. ## Citation ```bibtex @misc{lac2026neommesingletowermultimodalnativemultilingual, title={NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference}, author={Aurélien Lac and Tony Wu}, year={2026}, eprint={2609.01657}, archivePrefix={arXiv}, primaryClass={cs.IR}, url={https://arxiv.org/abs/2609.01657}, } ```