--- license: gemma base_model: google/embeddinggemma-300m base_model_relation: finetune pipeline_tag: sentence-similarity language: - en tags: - embeddinggemma - dense-retrieval - semantic-search - text-embeddings - document-retrieval - rag - fine-tuned - matryoshka - onnx - vulkan --- # EmbeddingGemma-300M memory retriever v2 An English embedding model for finding the passages that answer a question in notes, procedures, decision records and technical documentation. It is a fine-tune of [Google's EmbeddingGemma-300M](https://huggingface.co/google/embeddinggemma-300m), packaged as one ONNX graph with pooling and normalization built in, so it runs locally without PyTorch. | Stronger private retrieval | Less forgetting | Local deployment | |---|---|---| | nDCG@10 **0.4418 → 0.5621**; precision@10 **49.71% → 64.16%** | Recovers **more than half** of the first fine-tune's loss on FiQA and SciFact | **300M parameters**, ONNX, CPU / CUDA / Vulkan | The private comparison uses **925 queries with known useful evidence**, drawn from 970 queries over **82,719 passages**, with reviewed results through rank 50. Daecore accepted some general-retrieval loss for stronger evidence retrieval on its document workflow. The tables show both the gains and remaining public-benchmark gaps. Embed queries and passages with the same revision, and re-embed stored passages when switching to v2. - **Output:** a normalized 768-dimensional vector for each query or passage - **Input limits:** 128 tokens per query and 1,024 per passage, counting prefixes and special tokens - **Runtimes:** ONNX Runtime 1.24.4 on CPU and CUDA; Vulkan through the ONNX Runtime WebGPU plugin - **Size:** about 1.23 GB for the FP32 graph and its external weights - **License:** [Gemma Terms of Use](https://ai.google.dev/gemma/terms) ## Quick start Install the runtime dependencies and download the package. Keep `model.onnx`, `model.onnx.data` and the tokenizer from the same package together. ```sh pip install onnxruntime==1.24.4 tokenizers numpy huggingface_hub hf download Daecore/embeddinggemma-300m-memory-ft-v2 --local-dir downloaded-model ``` This example embeds a query and a passage on CPU: ```python from pathlib import Path import numpy as np import onnxruntime as ort from tokenizers import Tokenizer root = Path("downloaded-model") tokenizer = Tokenizer.from_file(str(root / "tokenizer.json")) session = ort.InferenceSession( str(root / "model.onnx"), providers=["CPUExecutionProvider"] ) def embed(text, *, query=False): prefix = "task: search result | query: " if query else "title: none | text: " tokenizer.enable_truncation(max_length=128 if query else 1024) encoded = tokenizer.encode(prefix + text) return session.run(["embeddings"], { "input_ids": np.array([encoded.ids], dtype=np.int64), "attention_mask": np.array([encoded.attention_mask], dtype=np.int64), })[0][0] query = embed("What must happen before a database migration?", query=True) passage = embed("Take a verified backup before applying the migration.") print(float(query @ passage)) ``` Higher cosine similarity means a closer match. Always use the query and passage prefixes shown above; the truncation lengths in the example are the tested input limits and include those prefixes. The default output is 768 dimensions. The smaller-width measurements below show the tradeoff when vector storage matters. ## Role in retrieval Gemma supplies the semantic candidates in Daecore's hybrid search. BM25 adds lexical matches for terms, names and identifiers; reciprocal-rank fusion merges the two rankings, and Ettin reranks up to 50 unique passages. A separate selector chooses the returned prefix. Chunking and candidate coverage therefore affect the final results alongside embedding quality. The package also works as a standalone embedder with other retrieval systems. Coverage at **50 candidates** measures what is available to a reranker. Hit@50 asks whether that pool contains any useful evidence; known-positive recall@50 measures how much of the judged useful evidence it contains. A reranker can improve the order but cannot recover a passage outside its pool. Gemma's dense top 50 describes a dense-only pipeline. Daecore's actual ceiling depends on the 50 candidates admitted after Gemma and BM25 are fused and deduplicated. ## Evaluation ### Gemma alone: dense retrieval on Daecore data Both models search the same **82,719 passages**, using FP32 exact dense retrieval, the same frozen query forms and matched 128/1,024-token input limits. BM25 and Ettin do not contribute. Each cell reads upstream → **Daecore v2**. **Reference:** 925 queries with known grade-2/3 evidence · **122,231 graded query–passage pairs** for those queries, within the shared 127,665-pair top-50 reference (September 2026). The same subset applies to both models, including queries where either model misses all useful passages. | Cutoff | Hit | Precision | Known-positive recall | nDCG | |---|---:|---:|---:|---:| | 1 | 56.00% → **67.89%** | 56.00% → **67.89%** | 1.44% → **1.79%** | 0.4621 → **0.5366** | | 3 | 76.54% → **83.24%** | 53.30% → **66.15%** | 3.94% → **5.03%** | 0.4466 → **0.5440** | | 5 | 83.35% → **88.86%** | 52.38% → **65.71%** | 6.17% → **8.06%** | 0.4398 → **0.5486** | | 10 | 88.97% → **92.43%** | 49.71% → **64.16%** | 11.18% → **15.07%** | 0.4418 → **0.5621** | | 20 | 93.08% → **94.92%** | 46.08% → **60.94%** | 19.61% → **27.56%** | 0.4580 → **0.5888** | | 50 | 97.73% → **99.03%** | 40.48% → **52.92%** | 41.12% → **57.92%** | 0.5145 → **0.6525** | ![Gemma v2 dense retrieval compared with upstream](figures/gemma-comparison.svg) Every measured top-50 position is graded or explicitly excluded as ungradable, with no missing-label ranges. All 925 queries remain eligible at every dense cutoff. At rank 50, 164 upstream positions and 167 v2 positions are excluded without backfill. Precision pools retained positions. The corpus is not exhaustively labeled; the other 45 source queries have no *known* useful passage, which does not prove none exists. At 50, v2 finds useful evidence for **99.03%** of these queries and retrieves **57.92%** of their known useful passages, versus 97.73% and 41.12% upstream. The capped hybrid pool below covers fewer queries at 50. Adding BM25 and capping the fused list can displace dense candidates, while reranking can improve early precision. Grades 2–3 count as useful. nDCG uses gains 0, 1, 3 and 7 against the same reviewed reference for both models. Rankings from query forms are combined with reciprocal-rank fusion, and identical passage text is deduplicated. The [evaluation companion](evaluation/README.md) includes all-query coverage, anonymized grades, identities and metric code. This is a reused development panel, not an untouched test of new workspaces. ### Gemma alone: public dense retrieval Five public retrieval datasets show how much general retrieval quality survives task adaptation. Scores are **dense-only nDCG@10** over full corpora, with matched preprocessing and official relevance judgments. | Dataset | Queries | Upstream Gemma | Daecore v2 | |---|---:|---:|---:| | SciFact | 300 | 0.7876 | 0.7783 | | FiQA | 648 | 0.4741 | 0.4468 | | NFCorpus | 323 | 0.3932 | 0.3890 | | SciDocs | 1,000 | 0.1945 | 0.1852 | | ArguAna | 1,406 | 0.6432 | 0.6259 | V2 is below upstream on all five datasets, most on FiQA (−0.0273). The first Daecore fine-tune was measured only on FiQA and SciFact, where it scored 0.4009 and 0.7679; v2 closes more than half of each of those gaps to upstream. None of these datasets supplied training examples, but they informed development, so they are not untouched tests. ### Smaller embeddings with Matryoshka The model was trained at 768, 512, 256 and 128 dimensions. To use a smaller width, keep the first dimensions and normalize the shortened vector again. Use the same width for queries and passages. Each quality column below is **v2 nDCG@10**, computed from exact rankings of cached FP32 vectors. The Daecore column covers the same 925-query dense panel with reviewed top-10 labels at each width, against the same **122,231 graded pairs** for these queries. The 768-wide results reproduce the corresponding full-width scores. | Dimensions | Bytes per FP32 vector | Daecore | SciFact | FiQA | NFCorpus | SciDocs | ArguAna | |---|---:|---:|---:|---:|---:|---:|---:| | 768 (default) | 3,072 | 0.5621 | 0.7783 | 0.4468 | 0.3890 | 0.1852 | 0.6259 | | 512 | 2,048 | 0.5648 | 0.7841 | 0.4374 | 0.3876 | 0.1793 | 0.6137 | | 256 | 1,024 | 0.5323 | 0.7757 | 0.4091 | 0.3660 | 0.1708 | 0.5987 | | 128 | 512 | 0.5069 | 0.7488 | 0.3668 | 0.3311 | 0.1484 | 0.5533 | At 512 dimensions, each raw FP32 vector uses one-third less storage, with slightly higher Daecore nDCG@10 in this measurement and lower scores on four of the five public panels. At 256, vector storage falls by two-thirds with a larger quality cost. These byte counts exclude index overhead and compression; shorter vectors do not reduce encoder inference cost. Daecore uses 768, and smaller widths have not been qualified through the full hybrid pipeline. ### Full pipeline: the effect of Gemma v2 Only the embedder changes: **upstream Gemma → Daecore Gemma v2**. BM25, reciprocal-rank fusion and the fine-tuned Ettin reranker stay fixed. Both pipelines search the same **82,719 passages**. Each cell reads before → **after**, so the change measures Gemma's contribution inside hybrid retrieval. **Reference:** the same **925 known-answerable queries**, frozen before this comparison · **122,783 graded query–passage pairs** · October 2026. Retrieval misses remain included. Both cards share the same fully fine-tuned endpoint. | Return depth | Hit | Precision | Known-positive recall | nDCG | |---|---:|---:|---:|---:| | 1 | 73.99% → **75.95%** | 73.99% → **75.95%** | 2.01% → **2.12%** | 0.6295 → **0.6527** | | 3 | 86.81% → **89.41%** | 71.50% → **75.11%** | 5.49% → **5.89%** | 0.6121 → **0.6424** | | 5 | 89.95% → **92.76%** | 70.10% → **73.67%** | 8.53% → **9.32%** | 0.6054 → **0.6373** | | 10 | 93.08% → **94.27%** | 65.20% → **70.49%** | 15.25% → **16.99%** | 0.5959 → **0.6377** | | 20 | 95.68% → **95.46%** | 56.86% → **64.80%** | 25.18% → **29.63%** | 0.5873 → **0.6489** | | 50 | 96.97% → **97.62%** | 35.45% → **42.40%** | 36.85% → **45.83%** | 0.5374 → **0.6125** | At three results, useful-passage precision changes from **71.50% to 75.11%**; nDCG@10 changes from **0.5959 to 0.6377**. At 50, the hybrid pool's known-positive recall changes from **36.85% to 45.83%**. This measures evidence available to Ettin after fusion; reranking cannot recover passages outside those 50 slots. The gain is not uniform: Hit@20 falls from 95.68% to 95.46%, a difference of two queries, while precision, known-positive recall and nDCG improve at that depth. All three variants use matched **PyTorch FP16** reranking. Depth 1 uses **919 shared judged queries** across the three pipeline variants; depths 3–50 use all 925. Abstentions are excluded inside the original cutoff without backfill. Grades 2–3 count as useful. Precision pools retained positions; recall counts known useful passages. The expanded reference and matched runtime distinguish this table from the earlier CUDA/Vulkan serving check. Daecore returns 3–20 results; these fixed cutoffs do not evaluate its selector. The [evaluation companion](evaluation/README.md) provides the paired records, methods and separate serving checks. The [pipeline overview](https://huggingface.co/Daecore) explains retrieval and selection. ## Training data and objective Private supervision adapts Gemma to **evidence retrieval from document chunks**. Several passages can be useful for one question, including partial evidence. | Training ingredient | Purpose | |---|---| | Mostly synthetic organizational documents, plus one person's largely AI-written project docs | Exercise project notes, procedures, decisions and technical material | | Frozen structural chunks with section context and table/code handling | Train on the passage units the retrieval workflow consumes | | Queries written by four models from three provider families, without a designated answer | Ask for evidence without reducing every question to one target passage | | Model-judged relevance grades | Distinguish decisive, partial, related and irrelevant evidence | The [chunking guide](evaluation/chunking.md#retrieval-model-data) explains the source data and methods. The corpus retains earlier parser outputs; Gemma applies its own tokenizer and limits after chunking. These shared generation and judging processes limit generalization beyond the tested data. Public supervision uses existing evidence annotations from [HotpotQA](https://hotpotqa.github.io/), [MultiDoc2Dial](https://github.com/doc2dial/multidoc2dial) and [FinQA](https://github.com/czyssrs/FinQA): multi-document questions, document-grounded dialogue, and textual or table evidence for numerical questions. The model learns to retrieve that evidence, not to perform FinQA's calculations; FinQA is unrelated to the FiQA evaluation above. The objective combines supervised retrieval, an upstream-similarity retention penalty that limits drift from upstream Gemma's similarity structure, and a small preference for grade-3 over grade-2 private positives. Known positives and related passages are masked so they do not act as ordinary in-batch negatives. The run completed 2,260 updates, and validation selected **update 1,130**. Counts are unique questions and query–passage pairs seen by that checkpoint: | Source | Available questions | Questions seen | Positive pairs seen | |---|---:|---:|---:| | Private Daecore | 2,165 | 2,165 | 73,130 | | HotpotQA | 90,272 | 45,140 | 90,280 | | MultiDoc2Dial | 4,931 | 2,466 | 2,508 | | FinQA | 5,754 | 2,905 | 4,966 | The checkpoint also saw 59,086 unique explicit private negative pairs; public groups used masked in-batch comparisons instead. Question and source weighting keep the largest pool from dominating through size alone.
Training recipe | Setting | Value | |---|---| | Initialization | Upstream Gemma; fresh optimizer | | Adapter | Attention LoRA, rank 16, alpha 32; merged for serving | | Effective batch | 64 question groups; up to four positives and four explicit negatives per group | | Source weights | Private 2/3; each public source 1/9 | | Optimizer | AdamW, learning rate 5e-5, weight decay 0.01 | | Schedule | 10% warmup, cosine decay over 2,260 updates | | Precision | FP16 with dynamic loss scaling | | Retention / graded preference weights | 2.0 / 0.25 | | Embedding widths and loss weights | 768 / 512 / 256 / 128, weighted 1 / 0.25 / 0.125 / 0.0625 | Small validation checks selected the checkpoint before full scoring. The recipe was run once; variation across training seeds is unmeasured.
## Runtime and limits The FP32 ONNX graph includes mean pooling, the learned projection and normalization. CPU, CUDA and Vulkan passed a 74-vector check covering token limits, mixed lengths and concurrent query and passage calls, with maximum differences from the Torch reference below 5.5e-7 and unchanged rankings on the tested inputs. Bounded mask operations were rewritten for Vulkan; learned weights are unchanged. Vulkan runs through the native ONNX Runtime WebGPU plugin, [`onnxruntime-ep-webgpu==0.4.0`](https://pypi.org/project/onnxruntime-ep-webgpu/), with `onnxruntime==1.24.4`. Register the plugin library, add its device to the session options and create the session without a providers list. Set `dawnBackendType` to `Vulkan`, because Windows can otherwise select Direct3D 12. The tested options also set `enableInt64=1`, `powerPreference=high-performance`, `validationMode=basic` and `storageBufferCacheMode=lazyRelease`. Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available. The task-specific data is mostly generated and model-judged, with no human reference panel, so it describes this development distribution rather than other users' documents. The dense and full-pipeline panels share queries and a pooled relevance reference, but evaluate different retrieval paths. Some questions omit the project or task context needed for an unambiguous judgment. The 970-query source panel uses natural-language questions and reports no separate terse-query or paraphrase robustness scores. Passage relevance does not measure distinct-fact coverage or downstream agent success. ## License [Gemma Terms of Use](https://ai.google.dev/gemma/terms) and the Gemma Prohibited Use Policy. The package includes the required terms, attribution and modification notice.