# EmbeddingGemma v2 evaluation These records support the [model card's](../README.md) results: Gemma's dense retrieval on private and public data, its smaller embedding widths, and the full retrieval pipeline. They contain scores and anonymized relevance grades, without private queries or document text. The [chunking guide](chunking.md#retrieval-model-data) explains the synthetic source documents, passage preparation and relevance supervision behind the private evaluations. | File | Contents | |---|---| | `retriever.json` | Upstream and v2 reviewed top-50 dense rankings for all 970 queries over 82,719 passages | | `dimensions.json` | V2 public nDCG@10 and private top-10 graded rankings at four embedding widths | | `promotion.json` | Ranked grades for the predecessor-to-v2 pipeline comparison and public nDCG@10 for upstream and v2 | | `pipeline.json` | Matched upstream-to-fine-tuned component changes on the frozen 925-query hybrid panel | | `serving.json` | Separate 970-query CUDA/Vulkan replay with the September reference | | `*-summary.json` | Summaries that `metrics.py` reproduces | | `serving-qualification.json` | CPU, CUDA and Vulkan serving checks for this model | | `metrics.py` | Metric code | | `figures.py`, `svg_figures.py` | Code to regenerate the Gemma comparison figure | ## Reproduce the tables From this directory, run with Python 3.11 or later: ```sh python metrics.py --directory . python figures.py --output ../figures ``` No packages or network access are needed. Each result matches its corresponding summary. The dense card table uses `retriever-summary.json`'s `answerable` section; the pipeline table uses `pipeline-summary.json`, comparing `upstream_gemma` with `finetuned`. `private.answerable` supplies the private width column. The dense and serving records retain all-query results in their outer `models` sections. This recomputes metrics from saved results; running private retrieval again would require the private corpus. ## Dense retrieval **Private data.** `retriever.json` compares upstream Gemma with v2 on all 970 queries over the same 82,719 passage texts. Both use FP32 inference, the same prefixes and 128-query/1,024-passage token limits, and 768 dimensions. Exact cosine rankings take the top 50 for each frozen query form, then merge them with reciprocal-rank fusion (constant 60). Identical text is deduplicated; there is no parent-document filter, BM25 or reranker. The record includes model identities, source hashes, top-50 grades and reference grade counts. The card reports the **925 queries with at least one known grade-2/3 passage** in the shared reference. This same subset is used for both models, including queries where a model retrieves no useful evidence. It contains **122,231 graded pairs** within the full 127,665-pair reference. The other 45 have no known useful passage; their true corpus-wide answerability is unresolved. This is a reused development panel, not a fresh holdout. The shared reference contains **127,665 graded query–passage pairs**. The top-50 extension added 40,231 grades and 222 abstentions from 40,453 previously unjudged pairs, covering both dense models, the hybrid pool and smaller-width top-10 results. Two independent GPT-6.1 Sol judges graded each pair; a fresh blind adjudicator resolved 7,765 ordinal disagreements. A semantic abstention from either primary judge remained excluded. The primary assistant read all 349 selected pairs: every abstention and final quotation flag, plus a stratified sample. Sixteen quotation repairs changed no grades. No person reviewed the labels. Every top-50 position is graded or explicitly excluded. All 925 answerable queries remain eligible at 1, 3, 5, 10, 20 and 50. Upstream excludes 164 positions at rank 50; v2 excludes 167. Cut the original prefix first, drop abstentions second, and never backfill. A wholly abstained prefix would omit the query from both arms at that depth. Precision pools retained positions. Recall counts known useful passages, not exhaustive corpus labels. nDCG uses gains 0, 1, 3 and 7 and reindexes retained positions. **Reference and population changes.** Expanding the reference preserved the old top-20 grades, rankings, Hit and precision. Across all 970 queries, dense nDCG@10 changed from 0.4506 to 0.4363 upstream and from 0.5748 to 0.5574 for v2 because newly judged relevant passages can raise ideal DCG. The earlier reference contained 87,434 graded pairs. The subsequent decision to headline the same 925 known-answerable queries changes the averaging population, not the labels or rankings. On that subset, nDCG@10 is **0.4418 upstream and 0.5621 for v2**. The September reference supports dense, smaller-width and provider results. The new component comparison adds its own reviewed overlay, described below. The historical predecessor comparison below keeps its original reference and population. **Public data.** The model card compares v2 with upstream EmbeddingGemma on SciFact, FiQA, NFCorpus, SciDocs and ArguAna. `promotion.json` includes per-query nDCG@10 for all five datasets, using full corpora, matched preprocessing and official relevance judgments. V2 remains below upstream on every panel. The previous fine-tune was measured only on FiQA (0.4009) and SciFact (0.7679); its other three results are absent, not estimated. These datasets supplied no training examples, but informed development. They are not untouched tests. **Matryoshka widths.** `dimensions.json` contains public per-query nDCG@10 and a `private` record with graded top-10 rankings at 768, 512, 256 and 128 dimensions. Vectors come from the selected checkpoint's cached FP32 outputs; smaller widths retain the first dimensions and are L2-normalized again. Public scoring uses exact cosine rankings, self-ID exclusion, stable corpus order for ties and official linear-gain nDCG. All 1,406 ArguAna queries remain, including five whose positives are absent from the corpus. These public scores are unchanged. Private scoring uses the same frozen query forms, dense-only fusion and expanded relevance reference as the 768-wide comparison. The same 925 answerable queries remain eligible at top 10; excluded positions are 22, 26, 16 and 37 respectively. The 768-wide results reproduce both references exactly. Smaller widths have not been qualified through the full pipeline and do not shorten the encoder's forward pass. ## Effect in the retrieval pipeline `pipeline.json` holds three matched arms. For this card, compare `upstream_gemma` with `finetuned`: only Gemma changes, while BM25, fusion and the fine-tuned Ettin remain fixed. The Ettin card uses the other baseline against the same fully fine-tuned endpoint. The comparison keeps the **925-query cohort frozen before scoring** and uses **122,783 graded query–passage pairs** for those queries. It adds 552 grades and 10 abstentions to their September reference of 122,231 pairs; old grades are unchanged. Two independent GPT-6.1 Sol judges and a fresh blind adjudicator resolved the new pairs, with 84 ordinal disagreements. The primary assistant reviewed 84 complete pairs, including every final quotation flag and abstention plus a stratified sample. These are model judgments, without a human reference panel. All three variants were scored anew in **PyTorch FP16**, using Transformers 5.2.0, Sentence Transformers 5.5.1, 128 query tokens and 1,024 passage tokens. Gemma uses cached FP32 embeddings at 768 dimensions with matched input limits. The actual first stage uses the frozen query forms, dense retrieval, BM25, reciprocal-rank fusion with constant 60, exact-text deduplication and 50 slots. Before substituting upstream Gemma's vectors, replay reproduced all 970 saved v2 candidate lists exactly. Ties preserve candidate order. Depth 1 uses **919 common judged queries across all three arms**; depths 3–50 retain all 925. Original prefixes are cut before abstentions are removed, with no backfill. Retrieval misses remain included. The reference supplies known-positive recall and ideal DCG; it is not exhaustive corpus labeling. Fixed cutoffs do not test the calibrated selector's choice of 3–20 results. The earlier ONNX provider replay retains its own runtime and September reference, so its absolute scores should not be substituted into this comparison. `serving.json` separately compares ONNX CUDA FP16 with Vulkan FP32. It keeps the September reference of 122,231 grades for the same 925 queries, with 920 eligible at top 1. Both providers use identical candidate pools. This checks serving paths, not the effect of fine-tuning. ### All-query coverage The September dense and serving records preserve all 970 source queries. This audit view includes the 45 without known useful evidence and is separate from the card's answerable-only quality table. Hybrid top 1 uses 965 queries after shared abstention exclusions; dense top 1 retains all 970. The counts below use the full 127,665-pair reference. | Retrieval path | Queries | Hit@3 | Hit@50 | nDCG@10 | |---|---:|---:|---:|---:| | Upstream Gemma, dense | 970 | 72.99% | 93.20% | 0.4363 | | Gemma v2, dense | 970 | 79.38% | 94.43% | 0.5574 | | Gemma v2 + BM25 + Ettin, CUDA | 970 | 85.36% | 93.09% | 0.6315 | A query with no known positive contributes zero Hit; grade-1 passages can still contribute to nDCG. These numbers should not be mixed with the card's 925-query results. ### Historical predecessor comparison This table comes from `promotion.json`, with its **original 75,488-pair relevance reference**. Its nDCG is not directly comparable with the expanded-reference table on the main card. Only the embedding model changed in this comparison. BM25, reciprocal-rank fusion and the [Ettin reranker](https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1) stayed fixed. The same score mapping selected returned prefixes in both arms. These results measure the whole pipeline on 970 reused queries over 82,719 passage texts, rather than Gemma alone. | Metric | Previous Gemma | Gemma v2 | |---|---:|---:| | Hit@1 (964 queries) | 72.61% | 72.30% | | Hit@3 | 83.92% | 85.36% | | Hit@5 | 85.88% | 88.45% | | Hit@10 | 87.84% | 89.90% | | Hit@20 | 90.10% | 91.03% | | nDCG@10 | 0.6764 | 0.6611 | | Selected-prefix precision | 80.41% | 80.34% | V2 found useful evidence for more queries, while the previous model placed higher-grade passages earlier. Seven reviewed grade corrections apply to both arms. Depths 3–20 retain all 970 queries; depth 1 omits the same six queries from both models. A shared set contains 65 unresolved query–passage judgments. Members of that set are excluded wherever they occur within an original cutoff, without filling their positions with lower-ranked passages. The record preserves those positions as `null` with an `excluded` flag. Grades 2 and 3 count as useful. Hit@k is the fraction of queries with useful evidence in the first k positions. Precision pools useful passages over retained positions. nDCG uses gain `2**grade - 1` and logarithmic rank discount; it reindexes retained passages and draws the ideal ranking from `reference_grade_counts`. Recall counts known useful passages, not distinct facts. Ties preserve candidate order. ## Query coverage The shared 970-query panel contains 620 generated-source questions, 200 questions written with potentially absent evidence, and 150 project-document questions. These are natural-language questions. In the pipeline replay, 483 use the original query alone; 487 multipart queries also use two deterministic subquestions. This exercises hybrid retrieval but does not provide a separately balanced terse-query or paraphrase benchmark. Training-query diversity is a separate claim from measured robustness on each query style. Dense retrieval can also serve keyword and identifier queries; “dense-only” identifies the retrieval method, not the query format. ## Serving checks and limits `serving-qualification.json` records this model's CPU, CUDA and Vulkan checks. The 74-vector panel covers token limits, mixed lengths and concurrent query and passage calls. It is a bounded serving check, separate from retrieval quality. The private documents and labels are mostly model-generated. Project-document sources come from one person's workspaces and are often AI-written. Some queries omit the project or subject needed for a clear relevance judgment. The panel was reused during development, training was run once, and there is no human reference panel. These results do not establish generalization across users, unique-fact coverage or downstream agent success.