sparbz commited on
Commit
2ad1174
·
verified ·
1 Parent(s): b3a7bba

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +58 -0
README.md ADDED
@@ -0,0 +1,58 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: ibm-granite/granite-embedding-311m-multilingual-r2
4
+ tags:
5
+ - gguf
6
+ - llama.cpp
7
+ - embeddings
8
+ - granite
9
+ - modernbert
10
+ - edge
11
+ ---
12
+
13
+ # granite-embedding-311m-multilingual-r2-GGUF
14
+
15
+ F16 GGUF conversion of [`ibm-granite/granite-embedding-311m-multilingual-r2`](https://huggingface.co/ibm-granite/granite-embedding-311m-multilingual-r2) for local serving with [llama.cpp](https://github.com/ggml-org/llama.cpp). Built and independently verified by ATF (Agent Taskflow) for edge-local embedding serving via `atf-serve`.
16
+
17
+ All model weights are © IBM, licensed Apache-2.0 (same as the base model).
18
+
19
+ ## Why this exists
20
+
21
+ IBM does not publish a GGUF for this model. This repo documents its own build end-to-end — source checksum, conversion command, and independent correctness verification — rather than asking you to trust an unverified re-hosted binary.
22
+
23
+ ## Conversion details
24
+
25
+ - **Source**: `ibm-granite/granite-embedding-311m-multilingual-r2`, `model.safetensors` (bf16, 623,341,952 bytes)
26
+ - **Tool**: `llama.cpp` built from source at commit `11924d4c17abc27383376a1ac6a24fa3e36c1c0c` (2026-08-02). This model's tokenizer (`granite-embed-multi-311m`, maps to `LLAMA_VOCAB_PRE_TYPE_GEMMA4`) is **not** recognized by llama.cpp release `b9204` or earlier — the registration landed upstream after that tag. A current build (or any release ≥ the commit that added it) is required both to *convert* and to *serve* this model; older binaries fail with `unknown pre-tokenizer type: 'granite-embed-multi-311m'` at load time, not at conversion time.
27
+ - **Command**:
28
+ ```
29
+ python3 convert_hf_to_gguf.py <model-dir> \
30
+ --outfile granite-embedding-311m-multilingual-r2-f16.gguf \
31
+ --outtype f16
32
+ ```
33
+ - **Output**: F16, 768-dim, 638,121,344 bytes.
34
+
35
+ ## Verification (independent, not vendor-claimed)
36
+
37
+ Embedded the same test sentence through both this GGUF (via `llama-server --embedding --pooling cls`) and the original HF model (via `sentence-transformers`, loaded directly from the source safetensors), then computed cosine similarity between the two output vectors.
38
+
39
+ | Check | Result |
40
+ |---|---|
41
+ | Output dimension | 768 (matches source `hidden_size`) |
42
+ | Cosine similarity vs. HF reference pipeline | **0.999970** |
43
+ | Required pooling mode | `cls` (matches source `classifier_pooling: "cls"` / `pooling_mode_cls_token: true` in `config.json`; mean pooling is **not** correct for this model) |
44
+
45
+ ## Usage
46
+
47
+ ```
48
+ llama-server --model granite-embedding-311m-multilingual-r2-f16.gguf \
49
+ --embedding --pooling cls --port 8089
50
+ ```
51
+
52
+ Requires a llama.cpp build that includes `granite-embed-multi-311m` tokenizer support (see Conversion details above — current upstream `master` has it; check your pinned release tag if serving fails with an `unknown pre-tokenizer type` error).
53
+
54
+ ```
55
+ curl http://127.0.0.1:8089/v1/embeddings \
56
+ -H "Content-Type: application/json" \
57
+ -d '{"input": "your text here", "model": "granite-embedding-311m"}'
58
+ ```