danbev commited on
Commit
27a93d4
·
verified ·
1 Parent(s): 3c62261

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +48 -0
README.md ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model:
3
+ - google/embeddinggemma-300M-qat-q4_0
4
+ ---
5
+ # embeddinggemma-300M-qat-q4_0 GGUF
6
+
7
+ Recommended way to run this model:
8
+
9
+ ```sh
10
+ llama-server -hf ggml-org/embeddinggemma-300M-qat-q4_0-GGUF
11
+ ```
12
+
13
+ Then the endpoint can be accessed at http://localhost:8080/embedding, for
14
+ example using `curl`:
15
+ ```console
16
+ curl --request POST \
17
+ --url http://localhost:8080/embedding \
18
+ --header "Content-Type: application/json" \
19
+ --data '{"input": "Hello embeddings"}' \
20
+ --silent
21
+ ```
22
+
23
+ Alternatively, the `llama-embedding` command line tool can be used:
24
+ ```sh
25
+ llama-embedding -hf ggml-org/embeddinggemma-300M-qat-q4_0-GGUF --verbose-prompt -p "Hello embeddings"
26
+ ```
27
+
28
+ #### embd_normalize
29
+ When a model uses pooling, or the pooling method is specified using `--pooling`,
30
+ the normalization can be controlled by the `embd_normalize` parameter.
31
+
32
+ The default value is `2` which means that the embeddings are normalized using
33
+ the Euclidean norm (L2). Other options are:
34
+ * -1 No normalization
35
+ * 0 Max absolute
36
+ * 1 Taxicab
37
+ * 2 Euclidean/L2
38
+ * \>2 P-Norm
39
+
40
+ This can be passed in the request body to `llama-server`, for example:
41
+ ```sh
42
+ --data '{"input": "Hello embeddings", "embd_normalize": -1}' \
43
+ ```
44
+
45
+ And for `llama-embedding`, by passing `--embd-normalize <value>`, for example:
46
+ ```sh
47
+ llama-embedding -hf ggml-org/embeddinggemma-300M-qat-q4_0-GGUF --embd-normalize -1 -p "Hello embeddings"
48
+ ```