ngquocvinh commited on
Commit
5dcf67b
·
verified ·
1 Parent(s): da064ac

Add Q3_K_S after text embedding smoke pass

Browse files
.gitattributes CHANGED
@@ -45,3 +45,4 @@ EmbeddingGemma-2-IQ1_M.gguf filter=lfs diff=lfs merge=lfs -text
45
  EmbeddingGemma-2-Q1_0.gguf filter=lfs diff=lfs merge=lfs -text
46
  EmbeddingGemma-2-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
47
  EmbeddingGemma-2-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
 
 
45
  EmbeddingGemma-2-Q1_0.gguf filter=lfs diff=lfs merge=lfs -text
46
  EmbeddingGemma-2-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
47
  EmbeddingGemma-2-Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
48
+ EmbeddingGemma-2-Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
EmbeddingGemma-2-Q3_K_S.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a6c72c2febc19a36afa8de82fb7e6cac8a45a2e7a20922612b2cc7c8b19270f6
3
+ size 142545536
SHA256SUMS.txt CHANGED
@@ -3,6 +3,7 @@ c4e59a215a94e4ab077e1a10c43df85ab1c63addc8fdf5515a5ee9b99c8a15f4 EmbeddingGemma
3
  4953a59ae0d8a725052407f458101ef81d8acb6a95f3293f4949ff3f594ef18a EmbeddingGemma-2-Q1_0.gguf
4
  67124cbba678cce4664bf9af72839a88ab12a9083821545c56847f41a54e96a1 EmbeddingGemma-2-Q2_K.gguf
5
  a5c902b15458c4886d6f1296bad71fd8084d7ac626cb8713778cf2af5e1b5a71 EmbeddingGemma-2-Q3_K_M.gguf
 
6
  ea25550b050fa8826707c5e44659a7c18259606dcc6ba449f771e3c75cf4e663 EmbeddingGemma-2-Q4_K_M.gguf
7
  4a9b92529b930be9f6cc858821b5d6ceea46c9631a05bd27f33a6942f46b63f6 EmbeddingGemma-2-Q4_K_S.gguf
8
  22f79226a02bfa2dd840414f342a004f93b3cb85e498923a5e089f03f6410b1f EmbeddingGemma-2-Q5_K_M.gguf
 
3
  4953a59ae0d8a725052407f458101ef81d8acb6a95f3293f4949ff3f594ef18a EmbeddingGemma-2-Q1_0.gguf
4
  67124cbba678cce4664bf9af72839a88ab12a9083821545c56847f41a54e96a1 EmbeddingGemma-2-Q2_K.gguf
5
  a5c902b15458c4886d6f1296bad71fd8084d7ac626cb8713778cf2af5e1b5a71 EmbeddingGemma-2-Q3_K_M.gguf
6
+ a6c72c2febc19a36afa8de82fb7e6cac8a45a2e7a20922612b2cc7c8b19270f6 EmbeddingGemma-2-Q3_K_S.gguf
7
  ea25550b050fa8826707c5e44659a7c18259606dcc6ba449f771e3c75cf4e663 EmbeddingGemma-2-Q4_K_M.gguf
8
  4a9b92529b930be9f6cc858821b5d6ceea46c9631a05bd27f33a6942f46b63f6 EmbeddingGemma-2-Q4_K_S.gguf
9
  22f79226a02bfa2dd840414f342a004f93b3cb85e498923a5e089f03f6410b1f EmbeddingGemma-2-Q5_K_M.gguf
reproducibility/artifact-measurements.tsv CHANGED
@@ -11,3 +11,4 @@ EmbeddingGemma-2-IQ1_M.gguf 103040640 0.103041 PASS_EMBEDDING_CPU
11
  EmbeddingGemma-2-Q1_0.gguf 66163328 0.066163 PASS_EMBEDDING_CPU
12
  EmbeddingGemma-2-Q5_K_S.gguf 210670208 0.210670 PASS_EMBEDDING_CPU
13
  EmbeddingGemma-2-Q4_K_S.gguf 178164352 0.178164 PASS_EMBEDDING_CPU
 
 
11
  EmbeddingGemma-2-Q1_0.gguf 66163328 0.066163 PASS_EMBEDDING_CPU
12
  EmbeddingGemma-2-Q5_K_S.gguf 210670208 0.210670 PASS_EMBEDDING_CPU
13
  EmbeddingGemma-2-Q4_K_S.gguf 178164352 0.178164 PASS_EMBEDDING_CPU
14
+ EmbeddingGemma-2-Q3_K_S.gguf 142545536 0.142546 PASS_EMBEDDING_CPU
reproducibility/manifest.md CHANGED
@@ -120,3 +120,11 @@ The fixed task evaluation uses the full MTEB STSBenchmark `test` split and the u
120
  - Quantization command: `llama-quantize --imatrix calibration/embeddinggemma-2/embeddinggemma-2-sts.imatrix.gguf embeddinggemma-2-text-BF16.gguf EmbeddingGemma-2-Q4_K_S.gguf Q4_K_S 8`.
121
  - CPU `llama-server --embedding --pooling mean -ngl 0` smoke passed with finite normalized 768-dimensional text vectors.
122
  - Build cgroup: `MemoryMax=32G`, `MemorySwapMax=0`; peak memory `926449664` bytes and swap peak `0` bytes.
 
 
 
 
 
 
 
 
 
120
  - Quantization command: `llama-quantize --imatrix calibration/embeddinggemma-2/embeddinggemma-2-sts.imatrix.gguf embeddinggemma-2-text-BF16.gguf EmbeddingGemma-2-Q4_K_S.gguf Q4_K_S 8`.
121
  - CPU `llama-server --embedding --pooling mean -ngl 0` smoke passed with finite normalized 768-dimensional text vectors.
122
  - Build cgroup: `MemoryMax=32G`, `MemorySwapMax=0`; peak memory `926449664` bytes and swap peak `0` bytes.
123
+ - Q4_K_S public Hub commit: `da064ac951ec91745594fe8699f3044842fc92ae`; remote LFS SHA256 and size match the local file.
124
+
125
+ ## Incremental experiment: Q3_K_S
126
+
127
+ - Artifact: `EmbeddingGemma-2-Q3_K_S.gguf`; SHA256 `a6c72c2febc19a36afa8de82fb7e6cac8a45a2e7a20922612b2cc7c8b19270f6`; decimal size `0.142546 GB`.
128
+ - Quantization command: `llama-quantize --imatrix calibration/embeddinggemma-2/embeddinggemma-2-sts.imatrix.gguf embeddinggemma-2-text-BF16.gguf EmbeddingGemma-2-Q3_K_S.gguf Q3_K_S 8`.
129
+ - CPU `llama-server --embedding --pooling mean -ngl 0` smoke passed with finite normalized 768-dimensional text vectors.
130
+ - Build cgroup: `MemoryMax=32G`, `MemorySwapMax=0`; peak memory `870584320` bytes and swap peak `0` bytes.