# Ettin reranker evaluation These records support the [model card's](../README.md) standalone and full-pipeline ranking results, plus separate CUDA/Vulkan serving checks. They contain anonymized relevance grades and scores, without private queries or passage text. The [chunking guide](chunking.md#retrieval-model-data) explains the synthetic source documents and prepared chunks that form the private candidate pools. | File | Contents | |---|---| | `reranker.json` | Grades and upstream and fine-tuned scores for 970 fixed pools of 50 passages | | `pipeline.json` | Matched component changes, including upstream and fine-tuned Ettin on Gemma v2 + BM25 candidates | | `serving.json` | Reviewed top-50 hybrid rankings from CUDA and Vulkan on the same 970 queries | | `*-summary.json` | The summaries that `metrics.py` reproduces | | `serving-qualification.json` | Serving checks, public reranking results and tested limits | | `metrics.py`, `figures.py`, `svg_figures.py` | Metric and figure code | ## Reproduce the tables From this directory, run with Python 3.11 or later: ```sh python metrics.py --directory . python figures.py --output ../figures ``` No packages or network access are needed. Each output matches its corresponding `*-summary.json`; the second command redraws the card's comparison figure. This recomputes saved results, rather than running inference again. The pipeline card table compares `upstream_ettin` with `finetuned` in `pipeline-summary.json`. The separate serving record preserves both answerable-only and all-query provider results. ## Ranking quality `reranker.json` holds the exact upstream and fine-tuned model identities and their rankings on the same 970 pools of 50 passages. All 48,500 pairs were judged; 873 pools contain useful evidence. Both models ran in PyTorch FP16 with the same 1,153-token pair construction, Transformers 5.2.0 and Sentence Transformers 5.5.1. Fresh inference on three complete pools checked the fine-tune's saved scores. This comparison measures the trained models, not ONNX latency. The model card presents results on the 873 answerable pools. The records also report all-query Hit and precision. Conditional nDCG and recall use the answerable pools. Grades 2 and 3 count as useful; nDCG uses gain `2**grade - 1` with logarithmic rank discount. Ties retain candidate order. The card shows cutoffs 1, 3, 5, 10 and 20, plus the full pool at 50. Daecore reranks up to 50 candidates and returns a selected prefix of 3–20. Recall here is relative to useful passages in the candidate pool. At 50, answerable-pool Hit and recall are 100% and precision is 42.68% for both models; these are properties of the shared pool. The 97 pools without useful evidence are excluded from the ranking-quality comparison because no ordering can make them answerable. The saved all-query metrics describe candidate coverage, not Ettin's conditional ranking quality. nDCG@50 remains sensitive to order: 0.7891 upstream and 0.8753 for the fine-tune. Random ordering is an exact within-pool reference: expected precision is R/N and Hit@k is `1 - C(N-R, k)/C(N, k)`, for N candidates and R useful passages. It already reaches 97.30% Hit@20 on answerable pools, so early precision and graded ranking are more informative here. The FiQA and SciFact results in `serving-qualification.json` rerank fixed pools of 50 candidates retrieved by upstream Gemma, with official relevance judgments. The fine-tune loses quality on FiQA and slightly improves SciFact. These are reranking results, not dense retrieval scores. ## Query coverage The shared 970-query panel contains 620 generated-source questions, 200 questions written with potentially absent evidence, and 150 project-document questions. These are natural-language questions. In the pipeline replay, 483 use the original query alone; 487 multipart queries also use two deterministic subquestions. This exercises hybrid retrieval but does not provide a separately balanced terse-query or paraphrase benchmark. Ettin trained on 294 answerable queries from an earlier lexical supplement. Its separate held-out artifacts are no longer available, so those historical checks cannot be recomputed and are not part of the comparison reported here. ## Effect in the retrieval pipeline `pipeline.json` compares upstream and fine-tuned Ettin on identical Gemma v2 + BM25 candidate pools (`upstream_ettin` → `finetuned`). A third arm changes Gemma alone for its companion card; both comparisons end at the same pipeline. At 50, Hit, precision and recall must match when only Ettin changes; nDCG still measures how well it orders that pool. The comparison keeps the **925-query cohort frozen before scoring** and uses **122,783 graded query–passage pairs** for those queries. It adds 552 grades and 10 abstentions to their September reference of 122,231 pairs; old grades are unchanged. Two independent GPT-6.1 Sol judges and a fresh blind adjudicator resolved the new pairs, with 84 ordinal disagreements. The primary assistant reviewed 84 complete pairs, including every final quotation flag and abstention plus a stratified sample. These are model judgments, without a human reference panel. All three variants were scored anew in **PyTorch FP16**, using Transformers 5.2.0, Sentence Transformers 5.5.1, 128 query tokens and 1,024 passage tokens. Gemma uses cached FP32 embeddings at 768 dimensions with matched input limits. The actual first stage uses the frozen query forms, dense retrieval, BM25, reciprocal-rank fusion with constant 60, exact-text deduplication and 50 slots. Before substituting upstream Gemma's vectors, replay reproduced all 970 saved v2 candidate lists exactly. Ties preserve candidate order. Depth 1 uses **919 common judged queries across all three arms**; depths 3–50 retain all 925. Original prefixes are cut before abstentions are removed, with no backfill. Retrieval misses remain included. The reference supplies known-positive recall and ideal DCG; it is not exhaustive corpus labeling. Fixed cutoffs do not test the calibrated selector's choice of 3–20 results. The earlier ONNX provider replay retains its own runtime and September reference, so its absolute scores should not be substituted into this comparison. ## CUDA and Vulkan The separate provider check retains the September reference: **122,231 graded pairs for 925 known-answerable queries**, within 127,665 pairs across all 970. Depth 1 uses 920 common judged queries; depths 3–50 use all 925. It does not compare upstream with fine-tuned Ettin. `serving.json` compares CUDA FP16 and Vulkan FP32 on 970 candidate pools retrieved with [Gemma v2](https://huggingface.co/Daecore/embeddinggemma-300m-memory-ft-v2) and BM25. Ettin's learned weights and the score mapping are unchanged. This is a retrieval-pipeline check on fixed inputs; it does not measure a change in Gemma. | Metric | CUDA FP16 | Vulkan FP32 | |---|---:|---:| | Hit@1 (920 queries) | 75.87% | 75.98% | | Hit@3 | 89.51% | 89.41% | | Hit@5 | 92.76% | 92.76% | | Hit@10 | 94.27% | 94.27% | | Hit@20 | 95.46% | 95.46% | | Hit@50 | 97.62% | 97.62% | | nDCG@10 | 0.6378 | 0.6381 | | Selected-prefix precision | 81.88% | 81.82% | The top-50 completion shares its expanded reference with Gemma's dense and smaller-width evaluations: 127,665 graded query–passage pairs. It adds 40,231 numeric grades and 222 abstentions from 40,453 pairs across those panels. Two independent GPT-6.1 Sol judges and a blind adjudicator handled disagreements; the primary assistant reviewed 349 complete pairs, including every abstention and final quotation flag. Sixteen quotation repairs changed no grades. These are model judgments without human review. Existing grades and rankings are unchanged. The earlier reference expansion changed recall and ideal DCG; this provider table conditions quality on the shared 925 known-answerable queries. Each provider omits 52 abstained positions at rank 20 and 93 at rank 50, with no backfill. Precision pools retained positions and nDCG reindexes them. The standalone fixed-pool Ettin comparison and public benchmark labels are unchanged. ### All-query coverage The September serving record also preserves all 970 source queries, including 45 with no known useful evidence. This separate audit view uses the full 127,665-pair reference; it does not replace the answerable-only card table. | Retrieval path | Queries | Hit@3 | Hit@50 | nDCG@10 | |---|---:|---:|---:|---:| | Gemma v2 + BM25 + Ettin, CUDA | 970 | 85.36% | 93.09% | 0.6315 | ### Serving latency On 20 matched pools replayed twice, second-pass Vulkan reranking took 1.56 s median and 2.11 s at the 95th percentile, against 1.92 s and 2.70 s for the previous DirectML graph. Measurements used one Windows x64 machine with an RTX 3060 Ti (8 GB). They do not establish performance on other GPUs or Linux. `serving-qualification.json` records the checks and their limits. ## Limits Some questions omit a project or subject, making relevance ambiguous. The private panel is mostly generated and model-judged, has no human reference, and was reused during development. Training-seed variation is unmeasured. A reranker cannot recover passages absent from its candidates, and several useful passages may repeat one fact. Passage-level scores do not establish unique-fact coverage, answer completeness or agent success.