Instructions to use atfai/granite-embedding-311m-multilingual-r2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use atfai/granite-embedding-311m-multilingual-r2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf atfai/granite-embedding-311m-multilingual-r2-GGUF:F16 # Run inference directly in the terminal: llama cli -hf atfai/granite-embedding-311m-multilingual-r2-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf atfai/granite-embedding-311m-multilingual-r2-GGUF:F16 # Run inference directly in the terminal: llama cli -hf atfai/granite-embedding-311m-multilingual-r2-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf atfai/granite-embedding-311m-multilingual-r2-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf atfai/granite-embedding-311m-multilingual-r2-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf atfai/granite-embedding-311m-multilingual-r2-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf atfai/granite-embedding-311m-multilingual-r2-GGUF:F16
Use Docker
docker model run hf.co/atfai/granite-embedding-311m-multilingual-r2-GGUF:F16
- LM Studio
- Jan
- Ollama
How to use atfai/granite-embedding-311m-multilingual-r2-GGUF with Ollama:
ollama run hf.co/atfai/granite-embedding-311m-multilingual-r2-GGUF:F16
- Unsloth Desktop
- Docker Model Runner
How to use atfai/granite-embedding-311m-multilingual-r2-GGUF with Docker Model Runner:
docker model run hf.co/atfai/granite-embedding-311m-multilingual-r2-GGUF:F16
- Lemonade
How to use atfai/granite-embedding-311m-multilingual-r2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull atfai/granite-embedding-311m-multilingual-r2-GGUF:F16
Run and chat with the model
lemonade run user.granite-embedding-311m-multilingual-r2-GGUF-F16
List all available models
lemonade list
- Atomic Chat
granite-embedding-311m-multilingual-r2-GGUF
F16 GGUF conversion of ibm-granite/granite-embedding-311m-multilingual-r2 for local serving with llama.cpp. Converted and independently verified by ATF (Agent Taskflow) for edge-local embedding serving via atf-serve.
This is a format conversion only โ no weights were modified, retrained, or fine-tuned. All model weights are ยฉ IBM, licensed Apache-2.0 (same as the base model). This repository is not affiliated with or endorsed by IBM.
Why this exists
IBM does not publish a GGUF for this model. This repo documents its own build end-to-end โ source checksum, conversion command, and independent correctness verification โ rather than asking you to trust an unverified re-hosted binary.
Conversion details
- Source:
ibm-granite/granite-embedding-311m-multilingual-r2,model.safetensors(bf16, 623,341,952 bytes) - Tool:
llama.cppbuilt from source at commit11924d4c17abc27383376a1ac6a24fa3e36c1c0c(2026-08-02). This model's tokenizer (granite-embed-multi-311m, maps toLLAMA_VOCAB_PRE_TYPE_GEMMA4) is not recognized by llama.cpp releaseb9204or earlier โ the registration landed upstream after that tag. A current build (or any release โฅ the commit that added it) is required both to convert and to serve this model; older binaries fail withunknown pre-tokenizer type: 'granite-embed-multi-311m'at load time, not at conversion time. - Command:
python3 convert_hf_to_gguf.py <model-dir> \ --outfile granite-embedding-311m-multilingual-r2-f16.gguf \ --outtype f16 - Output: F16, 768-dim, 638,121,344 bytes.
Verification (independent, not vendor-claimed)
Embedded the same test sentence through both this GGUF (via llama-server --embedding --pooling cls) and the original HF model (via sentence-transformers, loaded directly from the source safetensors), then computed cosine similarity between the two output vectors.
| Check | Result |
|---|---|
| Output dimension | 768 (matches source hidden_size) |
| Cosine similarity vs. HF reference pipeline | 0.999970 |
| Required pooling mode | cls (matches source classifier_pooling: "cls" / pooling_mode_cls_token: true in config.json; mean pooling is not correct for this model) |
Usage
llama-server --model granite-embedding-311m-multilingual-r2-f16.gguf \
--embedding --pooling cls --port 8089
Requires a llama.cpp build that includes granite-embed-multi-311m tokenizer support (see Conversion details above โ current upstream master has it; check your pinned release tag if serving fails with an unknown pre-tokenizer type error).
curl http://127.0.0.1:8089/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"input": "your text here", "model": "granite-embedding-311m"}'
Converted by ATF โ agent orchestration with edge-local model serving.
- Downloads last month
- 54
16-bit