Weidows commited on
Commit
9b09022
·
verified ·
1 Parent(s): b5b9791

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +49 -0
README.md CHANGED
@@ -37,3 +37,52 @@ Evaluation set: STS-B test (1,379 sentence pairs, human similarity 0-5). Baselin
37
  - Q8_0, Q6_K, Q5_K_M and IQ4_XS show negligible quality loss (|Δρ| < 0.05%, Emb Cosine > 0.985) and are safe drop-in replacements.
38
  - Q4_K_M (4.85 bpw) shows a small but visible drop (Δρ ≈ +0.60%, Emb Cosine 0.985) — notably worse than the equally-sized IQ4_XS, so prefer IQ4_XS or Q5_K_M over Q4_K_M when size is comparable.
39
  - IQ3_M (3.76 bpw) is the only variant with a clearly measurable drop (Δρ ≈ +0.89%, Emb Cosine 0.93); use only when storage is critical.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
37
  - Q8_0, Q6_K, Q5_K_M and IQ4_XS show negligible quality loss (|Δρ| < 0.05%, Emb Cosine > 0.985) and are safe drop-in replacements.
38
  - Q4_K_M (4.85 bpw) shows a small but visible drop (Δρ ≈ +0.60%, Emb Cosine 0.985) — notably worse than the equally-sized IQ4_XS, so prefer IQ4_XS or Q5_K_M over Q4_K_M when size is comparable.
39
  - IQ3_M (3.76 bpw) is the only variant with a clearly measurable drop (Δρ ≈ +0.89%, Emb Cosine 0.93); use only when storage is critical.
40
+
41
+ ## Usage (llama.cpp GGUF)
42
+
43
+ All files here are GGUF and run with [llama.cpp](https://github.com/ggml-org/llama.cpp). Replace the model file with the quant you downloaded. Use `-ngl 999` to offload layers to GPU (omit or `-ngl 0` for CPU-only).
44
+
45
+ ### Text embedding — command line
46
+
47
+ ```bash
48
+ llama-embedding \
49
+ -m WeMM-Embedding-2B-Q5_K_M.gguf \
50
+ -p "Represent the meaning of this sentence." \
51
+ --pooling last
52
+ ```
53
+
54
+ ### Text embedding — HTTP server
55
+
56
+ ```bash
57
+ llama-server \
58
+ -m WeMM-Embedding-2B-Q5_K_M.gguf \
59
+ --embedding \
60
+ -ngl 999 --host 0.0.0.0 --port 8080
61
+ ```
62
+
63
+ Then request embeddings via the OpenAI-compatible endpoint:
64
+
65
+ ```bash
66
+ curl http://localhost:8080/v1/embeddings \
67
+ -H "Content-Type: application/json" \
68
+ -d '{"input": "Represent the meaning of this sentence.", "model": "WeMM-Embedding-2B-Q5_K_M"}'
69
+ ```
70
+
71
+ ### Multimodal (image / video) — HTTP server
72
+
73
+ The visual projector (`mmproj-WeMM-Embedding-2B-BF16.gguf`) is required for image and video inputs:
74
+
75
+ ```bash
76
+ llama-server \
77
+ -m WeMM-Embedding-2B-Q5_K_M.gguf \
78
+ --mmproj mmproj-WeMM-Embedding-2B-BF16.gguf \
79
+ --embedding \
80
+ -ngl 999 --host 0.0.0.0 --port 8080
81
+ ```
82
+
83
+ Send image/video inside the chat content the same way as the base model (interleave `image`/`video` before `text`).
84
+
85
+ ### Notes
86
+
87
+ - Output is a 2048-dim L2-normalized vector; matryoshka truncation (e.g. `--embd-normalize` + slicing) follows the base model's `matryoshka_dimensions` [64, 128, 256, 512, 1024, 2048].
88
+ - Q8_0 / Q4_K_M / BF16 are mirrored from `DreamBlooms/WeMM-Embedding-2B-GGUF`; Q6_K / Q5_K_M / IQ4_XS / IQ3_M were produced for this repo with `llama-quantize` from the same BF16 master.