hanxiao's picture
add top links: EIS, ArXiv, Blog
2ceae2e verified
|
Raw History Blame Contribute Delete
3.26 kB
---
pipeline_tag: feature-extraction
tags:
- gguf
- embedding
- eurobert
- llama-cpp
- jina-embeddings-v5
language:
- multilingual
base_model: jinaai/jina-embeddings-v5-text-nano
base_model_relation: quantized
inference: false
license: cc-by-nc-4.0
library_name: llama.cpp
---
# jina-embeddings-v5-text-nano-clustering-GGUF
GGUF quantizations of [jina-embeddings-v5-text-nano-clustering](https://huggingface.co/jinaai/jina-embeddings-v5-text-nano) using llama.cpp. A 239M parameter multilingual embedding model quantized for efficient inference.
[Elastic Inference Service](https://www.elastic.co/docs/explore-analyze/elastic-inference/eis) | [ArXiv](https://arxiv.org/abs/2602.15547) | [Blog](https://jina.ai/news/jina-embeddings-v5-text-distilling-4b-quality-into-sub-1b-multilingual-embeddings)
> [!IMPORTANT]
> We highly recommend to first read [this blog post for more technical details and customized llama.cpp build](https://jina.ai/news/optimizing-ggufs-for-decoder-only-embedding-models).
## Overview
<p align="center">
<img src="https://jina-ai-gmbh.ghost.io/content/images/2026/02/v5_architecture_1771470917.png" alt="jina-embeddings-v5-text Architecture" width="600px">
</p>
`jina-embeddings-v5-text-nano-clustering` is a task-specific embedding model for **clustering**, part of the [jina-embeddings-v5-text](https://huggingface.co/jinaai/jina-embeddings-v5-text-nano) model family.
| Feature | Value |
| --- | --- |
| Parameters | 239M |
| Task | `clustering` |
| Embedding Dimension | 768 |
| Matryoshka Dimensions | 32, 64, 128, 256, 512, 768 |
| Pooling Strategy | Last-token pooling |
| Base Model | [jina-embeddings-v5-text-nano](https://huggingface.co/jinaai/jina-embeddings-v5-text-nano) |
<p align="center">
<img src="https://jina-ai-gmbh.ghost.io/content/images/2026/02/v5_mmteb-4.png" alt="MMTEB Multilingual Benchmark" width="500px">
</p>
<p align="center">
<img src="https://jina-ai-gmbh.ghost.io/content/images/2026/02/v5_mteb_en-4.png" alt="MTEB English Benchmark" width="500px">
</p>
<p align="center">
<img src="https://jina-ai-gmbh.ghost.io/content/images/2026/02/v5_retrieval-4.png" alt="Retrieval Benchmark Results" width="500px">
</p>
## Usage with llama.cpp
<details open>
<summary>via <a href="https://www.elastic.co/docs/explore-analyze/elastic-inference/eis">Elastic Inference Service</a></summary>
The fastest way to use v5-text in production. Elastic Inference Service (EIS) provides managed embedding inference with built-in scaling, so you can generate embeddings directly within your Elastic deployment.
```bash
PUT _inference/text_embedding/jina-v5
{
"service": "elastic",
"service_settings": {
"model_id": "jina-embeddings-v5-text-nano"
}
}
```
See the [Elastic Inference Service documentation](https://www.elastic.co/docs/explore-analyze/elastic-inference/eis) for setup details.
</details>
```bash
# Build llama.cpp (upstream)
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp && cmake -B build && cmake --build build --config Release
# Run embedding
./build/bin/llama-embedding -m jina-embeddings-v5-text-nano-clustering-Q8_0.gguf \
--pooling last -p "Your text here"
```
## License
CC-BY-NC-4.0. For commercial use, please [contact us](https://jina.ai/contact-sales).