Instructions to use trithemius/granite-embedding-311m-multilingual-r2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use trithemius/granite-embedding-311m-multilingual-r2-GGUF with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("trithemius/granite-embedding-311m-multilingual-r2-GGUF") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use trithemius/granite-embedding-311m-multilingual-r2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16 # Run inference directly in the terminal: llama cli -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16 # Run inference directly in the terminal: llama cli -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
Use Docker
docker model run hf.co/trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
- LM Studio
- Jan
- Ollama
How to use trithemius/granite-embedding-311m-multilingual-r2-GGUF with Ollama:
ollama run hf.co/trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
- Unsloth Desktop
- Docker Model Runner
How to use trithemius/granite-embedding-311m-multilingual-r2-GGUF with Docker Model Runner:
docker model run hf.co/trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
- Lemonade
How to use trithemius/granite-embedding-311m-multilingual-r2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull trithemius/granite-embedding-311m-multilingual-r2-GGUF:F16
Run and chat with the model
lemonade run user.granite-embedding-311m-multilingual-r2-GGUF-F16
List all available models
lemonade list
- Atomic Chat
Granite Embedding 311M Multilingual R2 - GGUF
This repository provides a GGUF conversion of the IBM Granite embedding model:
Base model: ibm-granite/granite-embedding-311m-multilingual-r2
Original model card: https://huggingface.co/ibm-granite/granite-embedding-311m-multilingual-r2
Model Description
Granite Embedding 311M Multilingual R2 is a multilingual embedding model developed by IBM for semantic search, retrieval, RAG, clustering, and similarity tasks.
Key features:
- Multilingual support
- 311M parameters
- 768-dimensional embeddings
- Up to 32k context length
- Optimized for retrieval and semantic similarity
- Compatible with llama.cpp through GGUF conversion
Files
| File | Precision | Recommended Use |
|---|---|---|
granite-embedding-311m-multilingual-r2-F16.gguf |
F16 | Maximum quality and accuracy |
Conversion Details
This GGUF file was generated using the official llama.cpp conversion tools.
Conversion command
python convert_hf_to_gguf.py \
granite-embedding-311m-multilingual-r2 \
--outfile granite-embedding-311m-multilingual-r2-F16.gguf \
--outtype f16
llama.cpp version
Commit: 96fbe0039337a999613a983d66e2bfcc4bb554d7
Usage with llama.cpp
Embedding generation
llama-embedding \
-m granite-embedding-311m-multilingual-r2-F16.gguf \
-p "Artificial intelligence is transforming software engineering."
OpenAI-compatible server
llama-server \
-m granite-embedding-311m-multilingual-r2-F16.gguf \
--embedding
Example request:
curl http://localhost:8080/v1/embeddings \
-H "Content-Type: application/json" \
-d '{
"input": "Hello world"
}'
Intended Uses
This model is suitable for:
- Retrieval-Augmented Generation (RAG)
- Semantic search
- Document retrieval
- Similarity search
- Clustering
- Deduplication
- Cross-lingual retrieval
- Recommendation systems
Notes
This repository only provides a GGUF conversion of the original IBM model. All credit for the model architecture, training, and evaluation belongs to IBM Research.
Please refer to the original model card for:
- Training details
- Evaluation results
- Benchmark scores
- Limitations
- Responsible AI considerations
Original repository:
https://huggingface.co/ibm-granite/granite-embedding-311m-multilingual-r2
License
This GGUF conversion is distributed under the same license as the original model:
Apache License 2.0
Please verify license compatibility with your intended use case.
Acknowledgements
- IBM Research for developing the Granite Embedding model.
- The llama.cpp project for GGUF support and inference.
- The Hugging Face community for model hosting and distribution.
- Downloads last month
- 44
16-bit