Instructions to use Lorelum/granite-embedding-97m-multilingual-r2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Lorelum/granite-embedding-97m-multilingual-r2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Lorelum/granite-embedding-97m-multilingual-r2-GGUF:Q4_0 # Run inference directly in the terminal: llama cli -hf Lorelum/granite-embedding-97m-multilingual-r2-GGUF:Q4_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Lorelum/granite-embedding-97m-multilingual-r2-GGUF:Q4_0 # Run inference directly in the terminal: llama cli -hf Lorelum/granite-embedding-97m-multilingual-r2-GGUF:Q4_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Lorelum/granite-embedding-97m-multilingual-r2-GGUF:Q4_0 # Run inference directly in the terminal: ./llama-cli -hf Lorelum/granite-embedding-97m-multilingual-r2-GGUF:Q4_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Lorelum/granite-embedding-97m-multilingual-r2-GGUF:Q4_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Lorelum/granite-embedding-97m-multilingual-r2-GGUF:Q4_0
Use Docker
docker model run hf.co/Lorelum/granite-embedding-97m-multilingual-r2-GGUF:Q4_0
- LM Studio
- Jan
- Ollama
How to use Lorelum/granite-embedding-97m-multilingual-r2-GGUF with Ollama:
ollama run hf.co/Lorelum/granite-embedding-97m-multilingual-r2-GGUF:Q4_0
- Unsloth Desktop
- Docker Model Runner
How to use Lorelum/granite-embedding-97m-multilingual-r2-GGUF with Docker Model Runner:
docker model run hf.co/Lorelum/granite-embedding-97m-multilingual-r2-GGUF:Q4_0
- Lemonade
How to use Lorelum/granite-embedding-97m-multilingual-r2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Lorelum/granite-embedding-97m-multilingual-r2-GGUF:Q4_0
Run and chat with the model
lemonade run user.granite-embedding-97m-multilingual-r2-GGUF-Q4_0
List all available models
lemonade list
- Atomic Chat
Granite Embedding 97M Multilingual R2 โ Q4_0 GGUF
Q4_0 quantization of IBM Granite Embedding 97M Multilingual R2, prepared for Lorelum's local CPU embedding service. This is a derivative artifact maintained by Lorelum, not an official IBM release.
File
- Filename:
granite-q4_0.gguf - Size: 66,345,216 bytes
- SHA-256:
18e8ce8ce834790618e90d26bed465cca87362076f3042eb0d8eee0732596f59 - Embedding dimensions: 384
- Lorelum uses CLS pooling and L2 normalization, with no query or document prefix.
Provenance and quantization
Original model: ibm-granite/granite-embedding-97m-multilingual-r2.
F16 GGUF conversion by ATF:
- Source file:
granite-embedding-97m-multilingual-r2-f16.gguf - Source size: 206,403,072 bytes
- Source SHA-256:
74075aeea7bd9ac4e5d74755e216fa487f5b6cce9ac6c5a363936370585b1a38
Quantized from that F16 file, without requantization, using llama.cpp b10901, commit 28ff0958291ce3465fabd7bd679d4b0edd742bd9:
llama-quantize --pure --token-embedding-type q4_0 \
granite-embedding-97m-multilingual-r2-f16.gguf \
granite-q4_0.gguf Q4_0 4
Matrix weights and the token embedding table use Q4_0; normalization vectors remain F32. This is the pure Q4_0 artifact, not Q4_K_M.
Usage
Download with the Hugging Face CLI:
hf download Lorelum/granite-embedding-97m-multilingual-r2-GGUF \
granite-q4_0.gguf --local-dir ./models
Load it with a compatible llama.cpp embedding runtime. Lorelum validates the artifact size and SHA-256 before loading it. This artifact was tested with CPU inference on macOS arm64; publication alone does not establish compatibility or performance on other hardware.
License and attribution
The original IBM model and ATF conversion declare Apache-2.0. This repository includes the Apache-2.0 LICENSE, the upstream model card, and the source GGUF model card to retain source documentation and attribution. The modification made here is Q4_0 quantization; no additional training was performed.
- Downloads last month
- 229
4-bit