File size: 16,894 Bytes
bf5a9d3 559fbd8 bf5a9d3 559fbd8 6ca39f1 bf5a9d3 559fbd8 bf5a9d3 559fbd8 bf5a9d3 13060d2 d02f909 13060d2 d02f909 6ca39f1 559fbd8 bf5a9d3 559fbd8 bf5a9d3 6ca39f1 bf5a9d3 31e2941 13060d2 bf5a9d3 6ca39f1 d02f909 6ca39f1 d02f909 13060d2 cc80fc0 d02f909 6ca39f1 d02f909 13060d2 d02f909 766fddf 6ca39f1 cc80fc0 d02f909 6ca39f1 bf5a9d3 559fbd8 bf5a9d3 559fbd8 bf5a9d3 6ca39f1 13060d2 d02f909 13060d2 d02f909 13060d2 6ca39f1 766fddf 6ca39f1 766fddf 13060d2 766fddf 6ca39f1 766fddf bf5a9d3 13060d2 bf5a9d3 13060d2 31e2941 bf5a9d3 559fbd8 bf5a9d3 559fbd8 bf5a9d3 d02f909 bf5a9d3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 | ---
license: gemma
base_model: google/embeddinggemma-300m
base_model_relation: finetune
pipeline_tag: sentence-similarity
language:
- en
tags:
- embeddinggemma
- dense-retrieval
- semantic-search
- text-embeddings
- document-retrieval
- rag
- fine-tuned
- matryoshka
- onnx
- vulkan
---
# EmbeddingGemma-300M memory retriever v2
An English embedding model for finding the passages that answer a question in notes, procedures, decision records and technical documentation. It is a fine-tune of [Google's EmbeddingGemma-300M](https://huggingface.co/google/embeddinggemma-300m), packaged as one ONNX graph with pooling and normalization built in, so it runs locally without PyTorch.
| Stronger private retrieval | Less forgetting | Local deployment |
|---|---|---|
| nDCG@10 **0.4418 β 0.5621**; precision@10 **49.71% β 64.16%** | Recovers **more than half** of the first fine-tune's loss on FiQA and SciFact | **300M parameters**, ONNX, CPU / CUDA / Vulkan |
The private comparison uses **925 queries with known useful evidence**, drawn
from 970 queries over **82,719 passages**, with reviewed results through rank
50. Daecore accepted some general-retrieval loss for stronger evidence retrieval
on its document workflow. The tables show both the gains and remaining
public-benchmark gaps.
Embed queries and passages with the same revision, and re-embed stored passages when switching to v2.
- **Output:** a normalized 768-dimensional vector for each query or passage
- **Input limits:** 128 tokens per query and 1,024 per passage, counting prefixes and special tokens
- **Runtimes:** ONNX Runtime 1.24.4 on CPU and CUDA; Vulkan through the ONNX Runtime WebGPU plugin
- **Size:** about 1.23 GB for the FP32 graph and its external weights
- **License:** [Gemma Terms of Use](https://ai.google.dev/gemma/terms)
## Quick start
Install the runtime dependencies and download the package. Keep `model.onnx`, `model.onnx.data` and the tokenizer from the same package together.
```sh
pip install onnxruntime==1.24.4 tokenizers numpy huggingface_hub
hf download Daecore/embeddinggemma-300m-memory-ft-v2 --local-dir downloaded-model
```
This example embeds a query and a passage on CPU:
```python
from pathlib import Path
import numpy as np
import onnxruntime as ort
from tokenizers import Tokenizer
root = Path("downloaded-model")
tokenizer = Tokenizer.from_file(str(root / "tokenizer.json"))
session = ort.InferenceSession(
str(root / "model.onnx"), providers=["CPUExecutionProvider"]
)
def embed(text, *, query=False):
prefix = "task: search result | query: " if query else "title: none | text: "
tokenizer.enable_truncation(max_length=128 if query else 1024)
encoded = tokenizer.encode(prefix + text)
return session.run(["embeddings"], {
"input_ids": np.array([encoded.ids], dtype=np.int64),
"attention_mask": np.array([encoded.attention_mask], dtype=np.int64),
})[0][0]
query = embed("What must happen before a database migration?", query=True)
passage = embed("Take a verified backup before applying the migration.")
print(float(query @ passage))
```
Higher cosine similarity means a closer match. Always use the query and passage prefixes shown above; the truncation lengths in the example are the tested input limits and include those prefixes. The default output is 768 dimensions. The smaller-width measurements below show the tradeoff when vector storage matters.
## Role in retrieval
Gemma supplies the semantic candidates in Daecore's hybrid search. BM25 adds
lexical matches for terms, names and identifiers; reciprocal-rank fusion merges
the two rankings, and Ettin reranks up to 50 unique passages. A separate selector
chooses the returned prefix. Chunking and candidate coverage therefore affect
the final results alongside embedding quality. The package also works as a
standalone embedder with other retrieval systems.
Coverage at **50 candidates** measures what is available to a reranker. Hit@50
asks whether that pool contains any useful evidence; known-positive recall@50
measures how much of the judged useful evidence it contains. A reranker can
improve the order but cannot recover a passage outside its pool. Gemma's dense
top 50 describes a dense-only pipeline. Daecore's actual ceiling depends on
the 50 candidates admitted after Gemma and BM25 are fused and deduplicated.
## Evaluation
### Gemma alone: dense retrieval on Daecore data
Both models search the same **82,719 passages**, using FP32 exact dense
retrieval, the same frozen query forms and matched 128/1,024-token input
limits. BM25 and Ettin do not contribute. Each cell reads upstream β
**Daecore v2**.
**Reference:** 925 queries with known grade-2/3 evidence Β· **122,231 graded
queryβpassage pairs** for those queries, within the shared 127,665-pair top-50
reference (September 2026). The same subset applies to both models, including
queries where either model misses all useful passages.
| Cutoff | Hit | Precision | Known-positive recall | nDCG |
|---|---:|---:|---:|---:|
| 1 | 56.00% β **67.89%** | 56.00% β **67.89%** | 1.44% β **1.79%** | 0.4621 β **0.5366** |
| 3 | 76.54% β **83.24%** | 53.30% β **66.15%** | 3.94% β **5.03%** | 0.4466 β **0.5440** |
| 5 | 83.35% β **88.86%** | 52.38% β **65.71%** | 6.17% β **8.06%** | 0.4398 β **0.5486** |
| 10 | 88.97% β **92.43%** | 49.71% β **64.16%** | 11.18% β **15.07%** | 0.4418 β **0.5621** |
| 20 | 93.08% β **94.92%** | 46.08% β **60.94%** | 19.61% β **27.56%** | 0.4580 β **0.5888** |
| 50 | 97.73% β **99.03%** | 40.48% β **52.92%** | 41.12% β **57.92%** | 0.5145 β **0.6525** |

Every measured top-50 position is graded or explicitly excluded as ungradable,
with no missing-label ranges. All 925 queries remain eligible at every dense
cutoff. At rank 50, 164 upstream positions and 167 v2 positions are excluded
without backfill. Precision pools retained positions. The corpus is not
exhaustively labeled; the other 45 source queries have no *known* useful
passage, which does not prove none exists.
At 50, v2 finds useful evidence for **99.03%** of these queries and retrieves
**57.92%** of their known useful passages, versus 97.73% and 41.12% upstream.
The capped hybrid pool below covers fewer queries at 50. Adding BM25 and
capping the fused list can displace dense candidates, while reranking can
improve early precision.
Grades 2β3 count as useful. nDCG uses gains 0, 1, 3 and 7 against the same
reviewed reference for both models. Rankings from query forms are combined
with reciprocal-rank fusion, and identical passage text is deduplicated. The
[evaluation companion](evaluation/README.md) includes all-query
coverage, anonymized grades, identities and metric code. This is a reused
development panel, not an untouched test of new workspaces.
### Gemma alone: public dense retrieval
Five public retrieval datasets show how much general retrieval quality survives task adaptation. Scores are **dense-only nDCG@10** over full corpora, with matched preprocessing and official relevance judgments.
| Dataset | Queries | Upstream Gemma | Daecore v2 |
|---|---:|---:|---:|
| SciFact | 300 | 0.7876 | 0.7783 |
| FiQA | 648 | 0.4741 | 0.4468 |
| NFCorpus | 323 | 0.3932 | 0.3890 |
| SciDocs | 1,000 | 0.1945 | 0.1852 |
| ArguAna | 1,406 | 0.6432 | 0.6259 |
V2 is below upstream on all five datasets, most on FiQA (β0.0273). The first Daecore fine-tune was measured only on FiQA and SciFact, where it scored 0.4009 and 0.7679; v2 closes more than half of each of those gaps to upstream. None of these datasets supplied training examples, but they informed development, so they are not untouched tests.
### Smaller embeddings with Matryoshka
The model was trained at 768, 512, 256 and 128 dimensions. To use a smaller
width, keep the first dimensions and normalize the shortened vector again.
Use the same width for queries and passages. Each quality column below is
**v2 nDCG@10**, computed from exact rankings of cached FP32 vectors. The Daecore
column covers the same 925-query dense panel with reviewed top-10 labels at
each width, against the same **122,231 graded pairs** for these queries. The 768-wide results
reproduce the corresponding full-width scores.
| Dimensions | Bytes per FP32 vector | Daecore | SciFact | FiQA | NFCorpus | SciDocs | ArguAna |
|---|---:|---:|---:|---:|---:|---:|---:|
| 768 (default) | 3,072 | 0.5621 | 0.7783 | 0.4468 | 0.3890 | 0.1852 | 0.6259 |
| 512 | 2,048 | 0.5648 | 0.7841 | 0.4374 | 0.3876 | 0.1793 | 0.6137 |
| 256 | 1,024 | 0.5323 | 0.7757 | 0.4091 | 0.3660 | 0.1708 | 0.5987 |
| 128 | 512 | 0.5069 | 0.7488 | 0.3668 | 0.3311 | 0.1484 | 0.5533 |
At 512 dimensions, each raw FP32 vector uses one-third less storage, with
slightly higher Daecore nDCG@10 in this measurement and lower scores on four
of the five public panels. At 256, vector storage falls by two-thirds with a
larger quality cost. These byte counts exclude index overhead and compression;
shorter vectors do not reduce encoder inference cost. Daecore uses 768, and
smaller widths have not been qualified through the full hybrid pipeline.
### Full pipeline: the effect of Gemma v2
Only the embedder changes: **upstream Gemma β Daecore Gemma v2**. BM25,
reciprocal-rank fusion and the fine-tuned Ettin reranker stay fixed. Both
pipelines search the same **82,719 passages**. Each cell reads before β
**after**, so the change measures Gemma's contribution inside hybrid retrieval.
**Reference:** the same **925 known-answerable queries**, frozen before this
comparison Β· **122,783 graded queryβpassage pairs** Β· October 2026. Retrieval
misses remain included. Both cards share the same fully fine-tuned endpoint.
| Return depth | Hit | Precision | Known-positive recall | nDCG |
|---|---:|---:|---:|---:|
| 1 | 73.99% β **75.95%** | 73.99% β **75.95%** | 2.01% β **2.12%** | 0.6295 β **0.6527** |
| 3 | 86.81% β **89.41%** | 71.50% β **75.11%** | 5.49% β **5.89%** | 0.6121 β **0.6424** |
| 5 | 89.95% β **92.76%** | 70.10% β **73.67%** | 8.53% β **9.32%** | 0.6054 β **0.6373** |
| 10 | 93.08% β **94.27%** | 65.20% β **70.49%** | 15.25% β **16.99%** | 0.5959 β **0.6377** |
| 20 | 95.68% β **95.46%** | 56.86% β **64.80%** | 25.18% β **29.63%** | 0.5873 β **0.6489** |
| 50 | 96.97% β **97.62%** | 35.45% β **42.40%** | 36.85% β **45.83%** | 0.5374 β **0.6125** |
At three results, useful-passage precision changes from **71.50% to
75.11%**; nDCG@10 changes from **0.5959 to 0.6377**.
At 50, the hybrid pool's known-positive recall changes from **36.85%
to 45.83%**. This measures evidence available to Ettin after
fusion; reranking cannot recover passages outside those 50 slots.
The gain is not uniform: Hit@20 falls from 95.68% to 95.46%, a difference of
two queries, while precision, known-positive recall and nDCG improve at that depth.
All three variants use matched **PyTorch FP16** reranking. Depth 1 uses **919 shared judged
queries** across the three pipeline variants; depths 3β50 use all 925.
Abstentions are excluded inside the original cutoff without backfill.
Grades 2β3 count as useful. Precision pools retained positions; recall counts
known useful passages.
The expanded reference and matched runtime distinguish this table from the
earlier CUDA/Vulkan serving check. Daecore returns 3β20 results; these fixed
cutoffs do not evaluate its selector.
The [evaluation companion](evaluation/README.md) provides
the paired records, methods and separate serving checks. The
[pipeline overview](https://huggingface.co/Daecore) explains retrieval and selection.
## Training data and objective
Private supervision adapts Gemma to **evidence retrieval from document chunks**.
Several passages can be useful for one question, including partial evidence.
| Training ingredient | Purpose |
|---|---|
| Mostly synthetic organizational documents, plus one person's largely AI-written project docs | Exercise project notes, procedures, decisions and technical material |
| Frozen structural chunks with section context and table/code handling | Train on the passage units the retrieval workflow consumes |
| Queries written by four models from three provider families, without a designated answer | Ask for evidence without reducing every question to one target passage |
| Model-judged relevance grades | Distinguish decisive, partial, related and irrelevant evidence |
The [chunking guide](evaluation/chunking.md#retrieval-model-data) explains
the source data and methods. The corpus retains earlier parser outputs;
Gemma applies its own tokenizer and limits after chunking. These shared
generation and judging processes limit generalization beyond the tested data.
Public supervision uses existing evidence annotations from [HotpotQA](https://hotpotqa.github.io/), [MultiDoc2Dial](https://github.com/doc2dial/multidoc2dial) and [FinQA](https://github.com/czyssrs/FinQA): multi-document questions, document-grounded dialogue, and textual or table evidence for numerical questions. The model learns to retrieve that evidence, not to perform FinQA's calculations; FinQA is unrelated to the FiQA evaluation above.
The objective combines supervised retrieval, an upstream-similarity retention penalty that limits drift from upstream Gemma's similarity structure, and a small preference for grade-3 over grade-2 private positives. Known positives and related passages are masked so they do not act as ordinary in-batch negatives.
The run completed 2,260 updates, and validation selected **update 1,130**. Counts are unique questions and queryβpassage pairs seen by that checkpoint:
| Source | Available questions | Questions seen | Positive pairs seen |
|---|---:|---:|---:|
| Private Daecore | 2,165 | 2,165 | 73,130 |
| HotpotQA | 90,272 | 45,140 | 90,280 |
| MultiDoc2Dial | 4,931 | 2,466 | 2,508 |
| FinQA | 5,754 | 2,905 | 4,966 |
The checkpoint also saw 59,086 unique explicit private negative pairs; public groups used masked in-batch comparisons instead. Question and source weighting keep the largest pool from dominating through size alone.
<details>
<summary>Training recipe</summary>
| Setting | Value |
|---|---|
| Initialization | Upstream Gemma; fresh optimizer |
| Adapter | Attention LoRA, rank 16, alpha 32; merged for serving |
| Effective batch | 64 question groups; up to four positives and four explicit negatives per group |
| Source weights | Private 2/3; each public source 1/9 |
| Optimizer | AdamW, learning rate 5e-5, weight decay 0.01 |
| Schedule | 10% warmup, cosine decay over 2,260 updates |
| Precision | FP16 with dynamic loss scaling |
| Retention / graded preference weights | 2.0 / 0.25 |
| Embedding widths and loss weights | 768 / 512 / 256 / 128, weighted 1 / 0.25 / 0.125 / 0.0625 |
Small validation checks selected the checkpoint before full scoring. The recipe was run once; variation across training seeds is unmeasured.
</details>
## Runtime and limits
The FP32 ONNX graph includes mean pooling, the learned projection and normalization. CPU, CUDA and Vulkan passed a 74-vector check covering token limits, mixed lengths and concurrent query and passage calls, with maximum differences from the Torch reference below 5.5e-7 and unchanged rankings on the tested inputs. Bounded mask operations were rewritten for Vulkan; learned weights are unchanged.
Vulkan runs through the native ONNX Runtime WebGPU plugin, [`onnxruntime-ep-webgpu==0.4.0`](https://pypi.org/project/onnxruntime-ep-webgpu/), with `onnxruntime==1.24.4`. Register the plugin library, add its device to the session options and create the session without a providers list. Set `dawnBackendType` to `Vulkan`, because Windows can otherwise select Direct3D 12. The tested options also set `enableInt64=1`, `powerPreference=high-performance`, `validationMode=basic` and `storageBufferCacheMode=lazyRelease`.
Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available.
The task-specific data is mostly generated and model-judged, with no human reference panel, so it describes this development distribution rather than other users' documents. The dense and full-pipeline panels share queries and a pooled relevance reference, but evaluate different retrieval paths. Some questions omit the project or task context needed for an unambiguous judgment. The 970-query source panel uses natural-language questions and reports no separate terse-query or paraphrase robustness scores. Passage relevance does not measure distinct-fact coverage or downstream agent success.
## License
[Gemma Terms of Use](https://ai.google.dev/gemma/terms) and the Gemma Prohibited Use Policy. The package includes the required terms, attribution and modification notice.
|