Feature Extraction
GGUF
multilingual
llama.cpp
embedding
eurobert
llama-cpp
jina-embeddings-v5
🇪🇺 Region: EU
Instructions to use jinaai/jina-embeddings-v5-text-nano-clustering-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use jinaai/jina-embeddings-v5-text-nano-clustering-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf jinaai/jina-embeddings-v5-text-nano-clustering-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf jinaai/jina-embeddings-v5-text-nano-clustering-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf jinaai/jina-embeddings-v5-text-nano-clustering-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf jinaai/jina-embeddings-v5-text-nano-clustering-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf jinaai/jina-embeddings-v5-text-nano-clustering-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf jinaai/jina-embeddings-v5-text-nano-clustering-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf jinaai/jina-embeddings-v5-text-nano-clustering-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf jinaai/jina-embeddings-v5-text-nano-clustering-GGUF:Q4_K_M
Use Docker
docker model run hf.co/jinaai/jina-embeddings-v5-text-nano-clustering-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use jinaai/jina-embeddings-v5-text-nano-clustering-GGUF with Ollama:
ollama run hf.co/jinaai/jina-embeddings-v5-text-nano-clustering-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use jinaai/jina-embeddings-v5-text-nano-clustering-GGUF with Docker Model Runner:
docker model run hf.co/jinaai/jina-embeddings-v5-text-nano-clustering-GGUF:Q4_K_M
- Lemonade
How to use jinaai/jina-embeddings-v5-text-nano-clustering-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull jinaai/jina-embeddings-v5-text-nano-clustering-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.jina-embeddings-v5-text-nano-clustering-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
|
Download README.md from jinaai/jina-embeddings-v5-text-nano-clustering-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 3.26 kB
-
https://huggingface.co/jinaai/jina-embeddings-v5-text-nano-clustering-GGUF/resolve/main/README.md
- Command line
-
hf download hf://jinaai/jina-embeddings-v5-text-nano-clustering-GGUF/README.md
-
curl -L -o README.md https://huggingface.co/jinaai/jina-embeddings-v5-text-nano-clustering-GGUF/resolve/main/README.md
3.26 kB
| pipeline_tag: feature-extraction | |
| tags: | |
| - gguf | |
| - embedding | |
| - eurobert | |
| - llama-cpp | |
| - jina-embeddings-v5 | |
| language: | |
| - multilingual | |
| base_model: jinaai/jina-embeddings-v5-text-nano | |
| base_model_relation: quantized | |
| inference: false | |
| license: cc-by-nc-4.0 | |
| library_name: llama.cpp | |
| # jina-embeddings-v5-text-nano-clustering-GGUF | |
| GGUF quantizations of [jina-embeddings-v5-text-nano-clustering](https://huggingface.co/jinaai/jina-embeddings-v5-text-nano) using llama.cpp. A 239M parameter multilingual embedding model quantized for efficient inference. | |
| [Elastic Inference Service](https://www.elastic.co/docs/explore-analyze/elastic-inference/eis) | [ArXiv](https://arxiv.org/abs/2602.15547) | [Blog](https://jina.ai/news/jina-embeddings-v5-text-distilling-4b-quality-into-sub-1b-multilingual-embeddings) | |
| > [!IMPORTANT] | |
| > We highly recommend to first read [this blog post for more technical details and customized llama.cpp build](https://jina.ai/news/optimizing-ggufs-for-decoder-only-embedding-models). | |
| ## Overview | |
| <p align="center"> | |
| <img src="https://jina-ai-gmbh.ghost.io/content/images/2026/02/v5_architecture_1771470917.png" alt="jina-embeddings-v5-text Architecture" width="600px"> | |
| </p> | |
| `jina-embeddings-v5-text-nano-clustering` is a task-specific embedding model for **clustering**, part of the [jina-embeddings-v5-text](https://huggingface.co/jinaai/jina-embeddings-v5-text-nano) model family. | |
| | Feature | Value | | |
| | --- | --- | | |
| | Parameters | 239M | | |
| | Task | `clustering` | | |
| | Embedding Dimension | 768 | | |
| | Matryoshka Dimensions | 32, 64, 128, 256, 512, 768 | | |
| | Pooling Strategy | Last-token pooling | | |
| | Base Model | [jina-embeddings-v5-text-nano](https://huggingface.co/jinaai/jina-embeddings-v5-text-nano) | | |
| <p align="center"> | |
| <img src="https://jina-ai-gmbh.ghost.io/content/images/2026/02/v5_mmteb-4.png" alt="MMTEB Multilingual Benchmark" width="500px"> | |
| </p> | |
| <p align="center"> | |
| <img src="https://jina-ai-gmbh.ghost.io/content/images/2026/02/v5_mteb_en-4.png" alt="MTEB English Benchmark" width="500px"> | |
| </p> | |
| <p align="center"> | |
| <img src="https://jina-ai-gmbh.ghost.io/content/images/2026/02/v5_retrieval-4.png" alt="Retrieval Benchmark Results" width="500px"> | |
| </p> | |
| ## Usage with llama.cpp | |
| <details open> | |
| <summary>via <a href="https://www.elastic.co/docs/explore-analyze/elastic-inference/eis">Elastic Inference Service</a></summary> | |
| The fastest way to use v5-text in production. Elastic Inference Service (EIS) provides managed embedding inference with built-in scaling, so you can generate embeddings directly within your Elastic deployment. | |
| ```bash | |
| PUT _inference/text_embedding/jina-v5 | |
| { | |
| "service": "elastic", | |
| "service_settings": { | |
| "model_id": "jina-embeddings-v5-text-nano" | |
| } | |
| } | |
| ``` | |
| See the [Elastic Inference Service documentation](https://www.elastic.co/docs/explore-analyze/elastic-inference/eis) for setup details. | |
| </details> | |
| ```bash | |
| # Build llama.cpp (upstream) | |
| git clone https://github.com/ggml-org/llama.cpp | |
| cd llama.cpp && cmake -B build && cmake --build build --config Release | |
| # Run embedding | |
| ./build/bin/llama-embedding -m jina-embeddings-v5-text-nano-clustering-Q8_0.gguf \ | |
| --pooling last -p "Your text here" | |
| ``` | |
| ## License | |
| CC-BY-NC-4.0. For commercial use, please [contact us](https://jina.ai/contact-sales). | |