trithemius commited on
Commit
efca743
·
verified ·
1 Parent(s): a60ad9d

Upload folder using huggingface_hub

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ granite-embedding-311m-multilingual-r2-F16.gguf filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,137 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - multilingual
4
+ license: apache-2.0
5
+ base_model:
6
+ - ibm-granite/granite-embedding-311m-multilingual-r2
7
+ library_name: llama.cpp
8
+ pipeline_tag: sentence-similarity
9
+ tags:
10
+ - gguf
11
+ - llama.cpp
12
+ - embeddings
13
+ - multilingual
14
+ - retrieval
15
+ - rag
16
+ - sentence-transformers
17
+ - granite
18
+ - modernbert
19
+ ---
20
+
21
+ # Granite Embedding 311M Multilingual R2 - GGUF
22
+
23
+ This repository provides a GGUF conversion of the IBM Granite embedding model:
24
+
25
+ **Base model:** `ibm-granite/granite-embedding-311m-multilingual-r2`
26
+
27
+ Original model card:
28
+ https://huggingface.co/ibm-granite/granite-embedding-311m-multilingual-r2
29
+
30
+ ## Model Description
31
+
32
+ Granite Embedding 311M Multilingual R2 is a multilingual embedding model developed by IBM for semantic search, retrieval, RAG, clustering, and similarity tasks.
33
+
34
+ Key features:
35
+
36
+ - Multilingual support
37
+ - 311M parameters
38
+ - 768-dimensional embeddings
39
+ - Up to 32k context length
40
+ - Optimized for retrieval and semantic similarity
41
+ - Compatible with llama.cpp through GGUF conversion
42
+
43
+ ## Files
44
+
45
+ | File | Precision | Recommended Use |
46
+ |------|-----------|-----------------|
47
+ | `granite-embedding-311m-multilingual-r2-F16.gguf` | F16 | Maximum quality and accuracy |
48
+
49
+ ## Conversion Details
50
+
51
+ This GGUF file was generated using the official `llama.cpp` conversion tools.
52
+
53
+ ### Conversion command
54
+
55
+ ```bash
56
+ python convert_hf_to_gguf.py \
57
+ granite-embedding-311m-multilingual-r2 \
58
+ --outfile granite-embedding-311m-multilingual-r2-F16.gguf \
59
+ --outtype f16
60
+ ```
61
+
62
+ ### llama.cpp version
63
+
64
+ ```text
65
+ Commit: 96fbe0039337a999613a983d66e2bfcc4bb554d7
66
+ ```
67
+
68
+ ## Usage with llama.cpp
69
+
70
+ ### Embedding generation
71
+
72
+ ```bash
73
+ llama-embedding \
74
+ -m granite-embedding-311m-multilingual-r2-F16.gguf \
75
+ -p "Artificial intelligence is transforming software engineering."
76
+ ```
77
+
78
+ ### OpenAI-compatible server
79
+
80
+ ```bash
81
+ llama-server \
82
+ -m granite-embedding-311m-multilingual-r2-F16.gguf \
83
+ --embedding
84
+ ```
85
+
86
+ Example request:
87
+
88
+ ```bash
89
+ curl http://localhost:8080/v1/embeddings \
90
+ -H "Content-Type: application/json" \
91
+ -d '{
92
+ "input": "Hello world"
93
+ }'
94
+ ```
95
+
96
+ ## Intended Uses
97
+
98
+ This model is suitable for:
99
+
100
+ - Retrieval-Augmented Generation (RAG)
101
+ - Semantic search
102
+ - Document retrieval
103
+ - Similarity search
104
+ - Clustering
105
+ - Deduplication
106
+ - Cross-lingual retrieval
107
+ - Recommendation systems
108
+
109
+ ## Notes
110
+
111
+ This repository only provides a GGUF conversion of the original IBM model. All credit for the model architecture, training, and evaluation belongs to IBM Research.
112
+
113
+ Please refer to the original model card for:
114
+
115
+ - Training details
116
+ - Evaluation results
117
+ - Benchmark scores
118
+ - Limitations
119
+ - Responsible AI considerations
120
+
121
+ Original repository:
122
+
123
+ https://huggingface.co/ibm-granite/granite-embedding-311m-multilingual-r2
124
+
125
+ ## License
126
+
127
+ This GGUF conversion is distributed under the same license as the original model:
128
+
129
+ **Apache License 2.0**
130
+
131
+ Please verify license compatibility with your intended use case.
132
+
133
+ ## Acknowledgements
134
+
135
+ - IBM Research for developing the Granite Embedding model.
136
+ - The llama.cpp project for GGUF support and inference.
137
+ - The Hugging Face community for model hosting and distribution.
granite-embedding-311m-multilingual-r2-F16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ec57b095530c2b3106c7112abc5e504b0b7588f072df69f5040f4f742342cae7
3
+ size 638121440