|
Download evaluation/README.md from Daecore/ettin-150m-memory-reranker-ft-v1: direct link, hf CLI and curl.
- Browser
- Download file 7.9 kB
-
https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1/resolve/main/evaluation/README.md
- Command line
-
hf download hf://Daecore/ettin-150m-memory-reranker-ft-v1/evaluation/README.md
-
curl -L -o README.md https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1/resolve/main/evaluation/README.md
7.9 kB
| # Ettin reranker evaluation | |
| These records support the [model card's](../README.md) ranking results and compare Ettin's CUDA and Vulkan serving paths. They contain anonymized relevance grades and scores, without private queries or passage text. | |
| The [chunking guide](chunking.md#retrieval-model-data) explains the synthetic | |
| source documents and prepared chunks that form the private candidate pools. | |
| | File | Contents | | |
| |---|---| | |
| | `reranker.json` | Grades and upstream and fine-tuned scores for 970 fixed pools of 50 passages | | |
| | `serving.json` | Reviewed top-50 hybrid rankings from CUDA and Vulkan on the same 970 queries | | |
| | `reranker-summary.json`, `serving-summary.json` | The summaries that `metrics.py` reproduces | | |
| | `serving-qualification.json` | Serving checks, public reranking results and tested limits | | |
| | `metrics.py`, `figures.py`, `svg_figures.py` | Metric and figure code | | |
| ## Reproduce the tables | |
| From this directory, run with Python 3.11 or later: | |
| ```sh | |
| python metrics.py --directory . | |
| python figures.py --output ../figures | |
| ``` | |
| No packages or network access are needed. The outputs match `reranker-summary.json` and `serving-summary.json`; the second command redraws the card's comparison figure. This recomputes saved results, rather than running inference again. The shared pipeline table uses `serving-summary.json`'s `answerable` section; its outer `models` section preserves all-query coverage. | |
| ## Ranking quality | |
| `reranker.json` holds the exact upstream and fine-tuned model identities and their rankings on the same 970 pools of 50 passages. All 48,500 pairs were judged; 873 pools contain useful evidence. Both models ran in PyTorch FP16 with the same 1,153-token pair construction, Transformers 5.2.0 and Sentence Transformers 5.5.1. Fresh inference on three complete pools checked the fine-tune's saved scores. This comparison measures the trained models, not ONNX latency. | |
| The model card presents results on the 873 answerable pools. The records also report all-query Hit and precision. Conditional nDCG and recall use the answerable pools. Grades 2 and 3 count as useful; nDCG uses gain `2**grade - 1` with logarithmic rank discount. Ties retain candidate order. | |
| The card shows cutoffs 1, 3, 5, 10 and 20, plus the full pool at 50. Daecore reranks up to 50 candidates and returns a selected prefix of 3–20. Recall here is relative to useful passages in the candidate pool. At 50, answerable-pool Hit and recall are 100% and precision is 42.68% for both models; these are properties of the shared pool. The 97 pools without useful evidence are excluded from the ranking-quality comparison because no ordering can make them answerable. The saved all-query metrics describe candidate coverage, not Ettin's conditional ranking quality. nDCG@50 remains sensitive to order: 0.7891 upstream and 0.8753 for the fine-tune. | |
| Random ordering is an exact within-pool reference: expected precision is R/N and Hit@k is `1 - C(N-R, k)/C(N, k)`, for N candidates and R useful passages. It already reaches 97.30% Hit@20 on answerable pools, so early precision and graded ranking are more informative here. | |
| The FiQA and SciFact results in `serving-qualification.json` rerank fixed pools of 50 candidates retrieved by upstream Gemma, with official relevance judgments. The fine-tune loses quality on FiQA and slightly improves SciFact. These are reranking results, not dense retrieval scores. | |
| ## Query coverage | |
| The shared 970-query panel contains 620 generated-source questions, 200 | |
| questions written with potentially absent evidence, and 150 project-document | |
| questions. These are natural-language questions. In the pipeline replay, 483 | |
| use the original query alone; 487 multipart queries also use two deterministic | |
| subquestions. This exercises hybrid retrieval but does not provide a separately | |
| balanced terse-query or paraphrase benchmark. Ettin trained on 294 answerable | |
| queries from an earlier lexical supplement. Its separate held-out artifacts | |
| are no longer available, so those historical checks cannot be recomputed and | |
| are not part of the comparison reported here. | |
| ## CUDA and Vulkan | |
| Both model cards share the CUDA fixed-depth table from `serving.json`, | |
| using **925 queries with known grade-2/3 evidence** in a shared reference | |
| across dense and hybrid retrieval. These queries have **122,231 graded | |
| query–passage pairs**, within a full reference of 127,665 pairs across 970 | |
| queries. Membership comes from the reference, so pipeline retrieval misses | |
| remain included. This is a different population and candidate set from | |
| Ettin's historical 873-answerable-pool table. | |
| Depths 3–50 use all 925 queries. Depth 1 uses the same 920 for both providers; | |
| five are omitted because a first result was ungradable. Fixed cutoffs do not | |
| evaluate the selector's choice of a 3–20-item prefix. | |
| At 50, the hybrid pool reaches **97.62% Hit and 45.98% known-positive recall** | |
| on the shared answerable subset. This is candidate coverage against the wider | |
| reference. Ettin's standalone fixed-pool Hit and recall instead reach 100% | |
| by construction on its answerable pools. | |
| `serving.json` compares CUDA FP16 and Vulkan FP32 on 970 candidate pools retrieved with [Gemma v2](https://huggingface.co/Daecore/embeddinggemma-300m-memory-ft-v2) and BM25. Ettin's learned weights and the score mapping are unchanged. This is a retrieval-pipeline check on fixed inputs; it does not measure a change in Gemma. | |
| | Metric | CUDA FP16 | Vulkan FP32 | | |
| |---|---:|---:| | |
| | Hit@1 (920 queries) | 75.87% | 75.98% | | |
| | Hit@3 | 89.51% | 89.41% | | |
| | Hit@5 | 92.76% | 92.76% | | |
| | Hit@10 | 94.27% | 94.27% | | |
| | Hit@20 | 95.46% | 95.46% | | |
| | Hit@50 | 97.62% | 97.62% | | |
| | nDCG@10 | 0.6378 | 0.6381 | | |
| | Selected-prefix precision | 81.88% | 81.82% | | |
| The top-50 completion shares its expanded reference with Gemma's dense and | |
| smaller-width evaluations: 127,665 graded query–passage pairs. It adds 40,231 numeric | |
| grades and 222 abstentions from 40,453 pairs across those panels. Two | |
| independent GPT-6.1 Sol judges and a blind adjudicator handled disagreements; | |
| the primary assistant reviewed 349 complete pairs, including every abstention | |
| and final quotation flag. Sixteen quotation repairs changed no grades. These | |
| are model judgments without human review. | |
| Existing grades and rankings are unchanged. The earlier reference expansion | |
| changed recall and ideal DCG; the card now also conditions quality on the | |
| shared 925 known-answerable queries. Each provider omits 52 abstained positions | |
| at rank 20 and 93 at rank 50, with no backfill. Precision pools retained | |
| positions and nDCG reindexes them. The standalone fixed-pool Ettin comparison | |
| and public benchmark labels are unchanged. | |
| ### All-query coverage | |
| The saved records also preserve all 970 source queries, including 45 with no | |
| known useful evidence. This separate audit view uses the full 127,665-pair | |
| reference; it does not replace the answerable-only card table. | |
| | Retrieval path | Queries | Hit@3 | Hit@50 | nDCG@10 | | |
| |---|---:|---:|---:|---:| | |
| | Gemma v2 + BM25 + Ettin, CUDA | 970 | 85.36% | 93.09% | 0.6315 | | |
| ### Serving latency | |
| On 20 matched pools replayed twice, second-pass Vulkan reranking took 1.56 s median and 2.11 s at the 95th percentile, against 1.92 s and 2.70 s for the previous DirectML graph. Measurements used one Windows x64 machine with an RTX 3060 Ti (8 GB). They do not establish performance on other GPUs or Linux. `serving-qualification.json` records the checks and their limits. | |
| ## Limits | |
| Some questions omit a project or subject, making relevance ambiguous. The private panel is mostly generated and model-judged, has no human reference, and was reused during development. Training-seed variation is unmeasured. A reranker cannot recover passages absent from its candidates, and several useful passages may repeat one fact. Passage-level scores do not establish unique-fact coverage, answer completeness or agent success. | |