Sentence Similarity
GGUF
sentence-transformers
multilingual
llama.cpp
embeddings
retrieval
rag
granite
modernbert
feature-extraction
Instructions to use trithemius/granite-embedding-311m-multilingual-r2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use trithemius/granite-embedding-311m-multilingual-r2-GGUF with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("trithemius/granite-embedding-311m-multilingual-r2-GGUF") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use trithemius/granite-embedding-311m-multilingual-r2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16 # Run inference directly in the terminal: llama cli -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16 # Run inference directly in the terminal: llama cli -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
Use Docker
docker model run hf.co/trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
- LM Studio
- Jan
- Ollama
How to use trithemius/granite-embedding-311m-multilingual-r2-GGUF with Ollama:
ollama run hf.co/trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
- Unsloth Desktop
- Docker Model Runner
How to use trithemius/granite-embedding-311m-multilingual-r2-GGUF with Docker Model Runner:
docker model run hf.co/trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
- Lemonade
How to use trithemius/granite-embedding-311m-multilingual-r2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
Run and chat with the model
lemonade run user.granite-embedding-311m-multilingual-r2-GGUF-F16
List all available models
lemonade list
- Atomic Chat
|
Download README.md from trithemius/granite-embedding-311m-multilingual-r2-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 3.09 kB
-
https://huggingface.co/trithemius/granite-embedding-311m-multilingual-r2-GGUF/resolve/main/README.md
- Command line
-
hf download hf://trithemius/granite-embedding-311m-multilingual-r2-GGUF/README.md
-
curl -L -o README.md https://huggingface.co/trithemius/granite-embedding-311m-multilingual-r2-GGUF/resolve/main/README.md
3.09 kB
| language: | |
| - multilingual | |
| license: apache-2.0 | |
| base_model: | |
| - ibm-granite/granite-embedding-311m-multilingual-r2 | |
| library_name: llama.cpp | |
| pipeline_tag: sentence-similarity | |
| tags: | |
| - gguf | |
| - llama.cpp | |
| - embeddings | |
| - multilingual | |
| - retrieval | |
| - rag | |
| - sentence-transformers | |
| - granite | |
| - modernbert | |
| # Granite Embedding 311M Multilingual R2 - GGUF | |
| This repository provides a GGUF conversion of the IBM Granite embedding model: | |
| **Base model:** `ibm-granite/granite-embedding-311m-multilingual-r2` | |
| Original model card: | |
| https://huggingface.co/ibm-granite/granite-embedding-311m-multilingual-r2 | |
| ## Model Description | |
| Granite Embedding 311M Multilingual R2 is a multilingual embedding model developed by IBM for semantic search, retrieval, RAG, clustering, and similarity tasks. | |
| Key features: | |
| - Multilingual support | |
| - 311M parameters | |
| - 768-dimensional embeddings | |
| - Up to 32k context length | |
| - Optimized for retrieval and semantic similarity | |
| - Compatible with llama.cpp through GGUF conversion | |
| ## Files | |
| | File | Precision | Recommended Use | | |
| |------|-----------|-----------------| | |
| | `granite-embedding-311m-multilingual-r2-F16.gguf` | F16 | Maximum quality and accuracy | | |
| ## Conversion Details | |
| This GGUF file was generated using the official `llama.cpp` conversion tools. | |
| ### Conversion command | |
| ```bash | |
| python convert_hf_to_gguf.py \ | |
| granite-embedding-311m-multilingual-r2 \ | |
| --outfile granite-embedding-311m-multilingual-r2-F16.gguf \ | |
| --outtype f16 | |
| ``` | |
| ### llama.cpp version | |
| ```text | |
| Commit: 96fbe0039337a999613a983d66e2bfcc4bb554d7 | |
| ``` | |
| ## Usage with llama.cpp | |
| ### Embedding generation | |
| ```bash | |
| llama-embedding \ | |
| -m granite-embedding-311m-multilingual-r2-F16.gguf \ | |
| -p "Artificial intelligence is transforming software engineering." | |
| ``` | |
| ### OpenAI-compatible server | |
| ```bash | |
| llama-server \ | |
| -m granite-embedding-311m-multilingual-r2-F16.gguf \ | |
| --embedding | |
| ``` | |
| Example request: | |
| ```bash | |
| curl http://localhost:8080/v1/embeddings \ | |
| -H "Content-Type: application/json" \ | |
| -d '{ | |
| "input": "Hello world" | |
| }' | |
| ``` | |
| ## Intended Uses | |
| This model is suitable for: | |
| - Retrieval-Augmented Generation (RAG) | |
| - Semantic search | |
| - Document retrieval | |
| - Similarity search | |
| - Clustering | |
| - Deduplication | |
| - Cross-lingual retrieval | |
| - Recommendation systems | |
| ## Notes | |
| This repository only provides a GGUF conversion of the original IBM model. All credit for the model architecture, training, and evaluation belongs to IBM Research. | |
| Please refer to the original model card for: | |
| - Training details | |
| - Evaluation results | |
| - Benchmark scores | |
| - Limitations | |
| - Responsible AI considerations | |
| Original repository: | |
| https://huggingface.co/ibm-granite/granite-embedding-311m-multilingual-r2 | |
| ## License | |
| This GGUF conversion is distributed under the same license as the original model: | |
| **Apache License 2.0** | |
| Please verify license compatibility with your intended use case. | |
| ## Acknowledgements | |
| - IBM Research for developing the Granite Embedding model. | |
| - The llama.cpp project for GGUF support and inference. | |
| - The Hugging Face community for model hosting and distribution. | |