Granite Embedding 311M Multilingual R2 - GGUF

This repository provides a GGUF conversion of the IBM Granite embedding model:

Base model: ibm-granite/granite-embedding-311m-multilingual-r2

Original model card: https://huggingface.co/ibm-granite/granite-embedding-311m-multilingual-r2

Model Description

Granite Embedding 311M Multilingual R2 is a multilingual embedding model developed by IBM for semantic search, retrieval, RAG, clustering, and similarity tasks.

Key features:

  • Multilingual support
  • 311M parameters
  • 768-dimensional embeddings
  • Up to 32k context length
  • Optimized for retrieval and semantic similarity
  • Compatible with llama.cpp through GGUF conversion

Files

File Precision Recommended Use
granite-embedding-311m-multilingual-r2-F16.gguf F16 Maximum quality and accuracy

Conversion Details

This GGUF file was generated using the official llama.cpp conversion tools.

Conversion command

python convert_hf_to_gguf.py \
    granite-embedding-311m-multilingual-r2 \
    --outfile granite-embedding-311m-multilingual-r2-F16.gguf \
    --outtype f16

llama.cpp version

Commit: 96fbe0039337a999613a983d66e2bfcc4bb554d7

Usage with llama.cpp

Embedding generation

llama-embedding \
    -m granite-embedding-311m-multilingual-r2-F16.gguf \
    -p "Artificial intelligence is transforming software engineering."

OpenAI-compatible server

llama-server \
    -m granite-embedding-311m-multilingual-r2-F16.gguf \
    --embedding

Example request:

curl http://localhost:8080/v1/embeddings \
    -H "Content-Type: application/json" \
    -d '{
        "input": "Hello world"
    }'

Intended Uses

This model is suitable for:

  • Retrieval-Augmented Generation (RAG)
  • Semantic search
  • Document retrieval
  • Similarity search
  • Clustering
  • Deduplication
  • Cross-lingual retrieval
  • Recommendation systems

Notes

This repository only provides a GGUF conversion of the original IBM model. All credit for the model architecture, training, and evaluation belongs to IBM Research.

Please refer to the original model card for:

  • Training details
  • Evaluation results
  • Benchmark scores
  • Limitations
  • Responsible AI considerations

Original repository:

https://huggingface.co/ibm-granite/granite-embedding-311m-multilingual-r2

License

This GGUF conversion is distributed under the same license as the original model:

Apache License 2.0

Please verify license compatibility with your intended use case.

Acknowledgements

  • IBM Research for developing the Granite Embedding model.
  • The llama.cpp project for GGUF support and inference.
  • The Hugging Face community for model hosting and distribution.
Downloads last month
44
GGUF
Model size
0.3B params
Architecture
modern-bert
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for trithemius/granite-embedding-311m-multilingual-r2-GGUF

Quantized
(16)
this model