EmbeddingGemma-300M memory retriever v2
An English embedding model for finding the passages that answer a question in notes, procedures, decision records and technical documentation. It is a fine-tune of Google's EmbeddingGemma-300M, packaged as one ONNX graph with pooling and normalization built in, so it runs locally without PyTorch.
Version 2 succeeds the first Daecore fine-tune with new weights trained to stay closer to upstream Gemma on general retrieval. It still scores below upstream on all five public datasets tested, but on the two where the first fine-tune was also measured, it recovers more than half of that model's loss. Embed queries and passages with the same revision, and re-embed stored passages when switching to v2.
- Output: a normalized 768-dimensional vector for each query or passage
- Input limits: 128 tokens per query and 1,024 per passage, counting prefixes and special tokens
- Runtimes: ONNX Runtime 1.24.4 on CPU and CUDA; Vulkan through the ONNX Runtime WebGPU plugin
- Size: about 1.23 GB for the FP32 graph and its external weights
- License: Gemma Terms of Use
Quick start
Install the runtime dependencies and download the package. Keep model.onnx, model.onnx.data and the tokenizer from the same package together.
pip install onnxruntime==1.24.4 tokenizers numpy huggingface_hub
hf download Daecore/embeddinggemma-300m-memory-ft-v2 --local-dir downloaded-model
This example embeds a query and a passage on CPU:
from pathlib import Path
import numpy as np
import onnxruntime as ort
from tokenizers import Tokenizer
root = Path("downloaded-model")
tokenizer = Tokenizer.from_file(str(root / "tokenizer.json"))
session = ort.InferenceSession(
str(root / "model.onnx"), providers=["CPUExecutionProvider"]
)
def embed(text, *, query=False):
prefix = "task: search result | query: " if query else "title: none | text: "
tokenizer.enable_truncation(max_length=128 if query else 1024)
encoded = tokenizer.encode(prefix + text)
return session.run(["embeddings"], {
"input_ids": np.array([encoded.ids], dtype=np.int64),
"attention_mask": np.array([encoded.attention_mask], dtype=np.int64),
})[0][0]
query = embed("What must happen before a database migration?", query=True)
passage = embed("Take a verified backup before applying the migration.")
print(float(query @ passage))
Higher cosine similarity means a closer match. Always use the query and passage prefixes shown above; the truncation lengths in the example are the tested input limits and include those prefixes. The qualified output is 768 dimensions; smaller Matryoshka widths need their own quality check.
Role in retrieval
Gemma supplies the semantic candidates in Daecore's hybrid search. BM25 adds lexical matches for terms, names and identifiers; reciprocal-rank fusion merges the two rankings, and Ettin reranks up to 50 unique passages. A separate selector chooses the returned prefix. Chunking and candidate coverage therefore affect the final results alongside embedding quality. The package also works as a standalone embedder with other retrieval systems.
Evaluation
Five public retrieval datasets show how much general retrieval quality survives task adaptation. Scores are dense-only nDCG@10 over full corpora, with matched preprocessing and official relevance judgments.
| Dataset | Queries | Upstream Gemma | Daecore v2 |
|---|---|---|---|
| SciFact | 300 | 0.7876 | 0.7783 |
| FiQA | 648 | 0.4741 | 0.4468 |
| NFCorpus | 323 | 0.3932 | 0.3890 |
| SciDocs | 1,000 | 0.1945 | 0.1852 |
| ArguAna | 1,406 | 0.6432 | 0.6259 |
V2 is below upstream on all five datasets, most on FiQA (−0.0273). The first Daecore fine-tune was measured only on FiQA and SciFact, where it scored 0.4009 and 0.7679; v2 closes more than half of each of those gaps to upstream. None of these datasets supplied training examples, but they informed development, so they are not untouched tests.
Inside Daecore's full search pipeline, with BM25, fusion and the Ettin reranker held fixed, v2 found a useful passage for more of the 970 development queries than the previous fine-tune at every cutoff from 3 to 20; Hit@10 rose from 87.84% to 89.90%. Graded ranking moved the other way: nDCG@10, which rewards putting the highest-grade passages first, fell from 0.6764 to 0.6611. The evaluation companion gives the full table with its judging rules and metric code; the retrieval pipeline overview describes the components.
Training data and objective
Private supervision teaches passage relevance for agent retrieval. A question can have several useful passages; decisive evidence is graded above useful but incomplete evidence, and similar but unhelpful passages are explicit negatives. The documents are generated organizational material plus one person's project documentation, much of it AI-written. Four models from three provider families wrote the queries without seeing a designated answer, and model judges graded relevance under a common rubric. Construction, chunking, query writing and judging follow shared processes, so the data does not stand in for other users' workspaces.
Private training passages were prepared with frozen versions of Daecore's structural chunker: headings and section context, cuts near paragraph or sentence boundaries, and handling for tables and code blocks. The corpus includes retained chunks from earlier parser versions; it was not regenerated for this release. Gemma applies its own tokenizer and input limits after passage preparation. Changing boundaries can change whether evidence stays together and how much the model sees, so these results are tied to the evaluated passages.
Public supervision uses existing evidence annotations from HotpotQA, MultiDoc2Dial and FinQA: multi-document questions, document-grounded dialogue, and textual or table evidence for numerical questions. The model learns to retrieve that evidence, not to perform FinQA's calculations; FinQA is unrelated to the FiQA evaluation above.
The objective combines supervised retrieval, an upstream-similarity retention penalty that limits drift from upstream Gemma's similarity structure, and a small preference for grade-3 over grade-2 private positives. Known positives and related passages are masked so they do not act as ordinary in-batch negatives.
The run completed 2,260 updates, and validation selected update 1,130. Counts are unique questions and query–passage pairs seen by that checkpoint:
| Source | Available questions | Questions seen | Positive pairs seen |
|---|---|---|---|
| Private Daecore | 2,165 | 2,165 | 73,130 |
| HotpotQA | 90,272 | 45,140 | 90,280 |
| MultiDoc2Dial | 4,931 | 2,466 | 2,508 |
| FinQA | 5,754 | 2,905 | 4,966 |
The checkpoint also saw 59,086 unique explicit private negative pairs; public groups used masked in-batch comparisons instead. Question and source weighting keep the largest pool from dominating through size alone.
Training recipe
| Setting | Value |
|---|---|
| Initialization | Upstream Gemma; fresh optimizer |
| Adapter | Attention LoRA, rank 16, alpha 32; merged for serving |
| Effective batch | 64 question groups; up to four positives and four explicit negatives per group |
| Source weights | Private 2/3; each public source 1/9 |
| Optimizer | AdamW, learning rate 5e-5, weight decay 0.01 |
| Schedule | 10% warmup, cosine decay over 2,260 updates |
| Precision | FP16 with dynamic loss scaling |
| Retention / graded preference weights | 2.0 / 0.25 |
| Embedding widths and loss weights | 768 / 512 / 256 / 128, weighted 1 / 0.25 / 0.125 / 0.0625 |
Small validation checks selected the checkpoint before full scoring. The recipe was run once; variation across training seeds is unmeasured.
Runtime and limits
The FP32 ONNX graph includes mean pooling, the learned projection and normalization. CPU, CUDA and Vulkan passed a 74-vector check covering token limits, mixed lengths and concurrent query and passage calls, with maximum differences from the Torch reference below 5.5e-7 and unchanged rankings on the tested inputs. Bounded mask operations were rewritten for Vulkan; learned weights are unchanged.
Vulkan runs through the native ONNX Runtime WebGPU plugin, onnxruntime-ep-webgpu==0.4.0, with onnxruntime==1.24.4. Register the plugin library, add its device to the session options and create the session without a providers list. Set dawnBackendType to Vulkan, because Windows can otherwise select Direct3D 12. The tested options also set enableInt64=1, powerPreference=high-performance, validationMode=basic and storageBufferCacheMode=lazyRelease.
Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available.
The task-specific data is mostly generated and model-judged, with no human reference panel, so it describes this development distribution rather than other users' documents. A matched dense-retrieval comparison between v2 and upstream on this private panel was not measured. The 970-query pipeline panel uses natural-language questions and reports no separate terse-query or paraphrase robustness scores. Passage relevance does not measure distinct-fact coverage or downstream agent success.
License
Gemma Terms of Use and the Gemma Prohibited Use Policy. The package includes the required terms, attribution and modification notice.
- Downloads last month
- -
Model tree for Daecore/embeddinggemma-300m-memory-ft-v2
Base model
google/embeddinggemma-300m