Download README.md from Daecore/ettin-150m-memory-reranker-ft-v1: direct link, hf CLI and curl.
- Browser
- Download file 13.4 kB
-
https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1/resolve/main/README.md
- Command line
-
hf download hf://Daecore/ettin-150m-memory-reranker-ft-v1/README.md
-
curl -L -o README.md https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1/resolve/main/README.md
license: apache-2.0
base_model: cross-encoder/ettin-reranker-150m-v1
base_model_relation: finetune
pipeline_tag: text-ranking
language:
- en
tags:
- ettin
- modernbert
- cross-encoder
- reranker
- fine-tuned
- onnx
- vulkan
Ettin-150M memory reranker
An English cross-encoder that reorders search results so the most useful passages come first. Give it a query and candidate passages from notes, procedures, decision records or technical documentation; it returns one relevance score per pair. Place it after a keyword or embedding search when the first few results matter most. It is a fine-tune of Ettin Reranker 150M and runs locally through ONNX Runtime.
| More useful early results | Stronger graded ranking | Local deployment |
|---|---|---|
| Precision@3 59.76% β 80.03% | nDCG@10 0.5519 β 0.7407 | 150M parameters, ONNX, CUDA / Vulkan |
Measured against upstream on 873 answerable pools of 50 passages, drawn from 970 pools with 48,500 graded queryβpassage pairs. Daecore accepted a FiQA loss for stronger ordering of the document chunks an agent reads first; SciFact improves slightly. Component and full-pipeline results are reported separately below.
This revision adds a Vulkan FP32 graph, model.vulkan.onnx; the learned weights and the CUDA FP16 graph are unchanged. It omits the DirectML graph, which remains in the previous revision.
- Query limit: 128 content tokens, excluding special tokens
- Passage limit: 1,022 content tokens, excluding special tokens
- Output: one relevance score per pair; higher scores rank first
- Runtimes: ONNX Runtime 1.24.4 with CUDA FP16 (
model.onnx) or Vulkan FP32 (model.vulkan.onnx); the FP32 graph was also checked for correctness on CPU - License: Apache-2.0
Quick start
Install the CUDA build of ONNX Runtime and download the package. Keep model.onnx, model.onnx.data and the tokenizer from the same package together.
pip install "onnxruntime-gpu[cuda,cudnn]==1.24.4" tokenizers numpy huggingface_hub
hf download Daecore/ettin-150m-memory-reranker-ft-v1 --local-dir downloaded-model
This CUDA example scores two passages:
import numpy as np
import onnxruntime as ort
from pathlib import Path
from tokenizers import Tokenizer
root = Path("downloaded-model") # The exact downloaded serving package.
tokenizer = Tokenizer.from_file(f"{root}/tokenizer.json")
tokenizer.no_padding()
tokenizer.no_truncation()
ort.preload_dlls()
session = ort.InferenceSession(
f"{root}/model.onnx", providers=[("CUDAExecutionProvider", {"use_tf32": "0"})]
)
if "CUDAExecutionProvider" not in session.get_providers():
raise RuntimeError("CUDA did not initialize; check the ONNX Runtime CUDA dependencies.")
query = "What must happen before a database migration?"
passages = [
"Take a verified backup before applying the migration.",
"The dashboard theme uses a blue background.",
]
query_tokens = tokenizer.encode(query, add_special_tokens=False)
if len(query_tokens.ids) > 128:
raise ValueError("Query exceeds the qualified 128-token content limit.")
scores = []
for passage in passages:
passage_tokens = tokenizer.encode(passage, add_special_tokens=False)
passage_tokens.truncate(1022)
pair = tokenizer.post_process(query_tokens, passage_tokens, add_special_tokens=True)
inputs = {
"input_ids": np.array([pair.ids], dtype=np.int64),
"attention_mask": np.array([pair.attention_mask], dtype=np.int64),
}
scores.append(float(session.run(None, inputs)[0].reshape(-1)[0]))
print(sorted(zip(passages, scores), key=lambda item: item[1], reverse=True))
Higher scores rank first; raw scores are not probabilities. The input envelope is 128 query-content tokens plus 1,022 passage-content tokens and three special tokens, a maximum of 1,153 tokens. The example rejects an oversized query and truncates the passage, matching the tested input policy. Without CUDA, use the FP32 graph model.vulkan.onnx with Vulkan or the CPU provider, as described under Runtime and limits.
Evaluation
Ettin alone: ranking a fixed candidate pool
Upstream and fine-tuned Ettin scored the same 970 pools of 50 candidates, with all 48,500 pairs judged. The table uses the 873 answerable pools that contain at least one useful passage, so it measures ordering after retrieval. βUpstreamβ is the exact starting checkpoint without Daecore fine-tuning. The headline panel uses natural-language questions; it does not report separate terse-query or paraphrase robustness scores.
Ettin scores up to 50 candidates; Daecore returns a selected prefix of 3β20. Each cell below reads upstream β Daecore fine-tune.
| Cutoff | Hit | Precision | Pool recall | nDCG |
|---|---|---|---|---|
| 1 | 61.74% β 81.44% | 61.74% β 81.44% | 3.68% β 5.23% | 0.5277 β 0.7182 |
| 3 | 83.16% β 92.55% | 59.76% β 80.03% | 10.03% β 14.28% | 0.5180 β 0.7198 |
| 5 | 89.46% β 94.04% | 58.58% β 77.94% | 15.80% β 22.34% | 0.5264 β 0.7212 |
| 10 | 93.81% β 96.91% | 56.74% β 74.18% | 29.07% β 39.95% | 0.5519 β 0.7407 |
| 20 | 98.17% β 98.97% | 54.46% β 67.34% | 53.78% β 67.88% | 0.6154 β 0.7869 |
| 50 (full pool) | 100.00% β 100.00% | 42.68% β 42.68% | 100.00% β 100.00% | 0.7891 β 0.8753 |
Grades 2 and 3 count as useful. Hit asks whether at least one useful passage appears, precision measures the useful share of returned positions, and nDCG rewards higher grades near the top. Pool recall measures the fraction of useful passages found within these 50 candidates, not across the full corpus.
At 50, both models return the entire answerable pool. Hit and pool recall therefore reach 100% by construction. The 97 pools without useful evidence are excluded from this ranking-quality comparison: no ordering could make them answerable. Hit, precision and recall at 50 cannot change through reranking the same candidates, while nDCG@50 still measures their order. These historical fixed pools are separate from the Gemma v2 pipeline replay below; retrieval coverage is measured separately from Ettin's ordering quality.
Many pools contain several useful passages, so random ordering already reaches 97.30% Hit@20; early precision and graded ranking are more informative than deep Hit. The comparison runs both models in PyTorch FP16 with matched tokenization on reused development inputs. The evaluation companion includes all-query results, model identities and metric code.
Ettin alone: public reranking
Both models rerank the same 50 candidates that upstream Gemma retrieved for each query, scored with official relevance judgments:
| Dataset | Queries | Upstream Ettin | Daecore fine-tune |
|---|---|---|---|
| FiQA | 648 | 0.4860 | 0.4527 |
| SciFact | 300 | 0.7487 | 0.7544 |
These are nDCG@10 scores for reranking fixed candidate pools, not dense Gemma scores. Fine-tuning lowered FiQA and slightly raised SciFact. Later continuation recipes did not recover FiQA while keeping Daecore ranking quality, so the original learned weights remain.
Full pipeline: the effect of fine-tuning Ettin
Only the reranker changes: upstream Ettin β Daecore Ettin. Gemma v2, BM25 and reciprocal-rank fusion supply the same 50 candidates from 82,719 passages. Each cell reads before β after, isolating what fine-tuning adds to the ordering of the actual hybrid candidates.
Reference: the same 925 known-answerable queries, frozen before this comparison Β· 122,783 graded queryβpassage pairs Β· October 2026. Retrieval misses remain included. Both cards share the same fully fine-tuned endpoint.
| Return depth | Hit | Precision | Known-positive recall | nDCG |
|---|---|---|---|---|
| 1 | 56.69% β 75.95% | 56.69% β 75.95% | 1.44% β 2.12% | 0.4732 β 0.6527 |
| 3 | 79.03% β 89.41% | 56.04% β 75.11% | 4.23% β 5.89% | 0.4728 β 0.6424 |
| 5 | 85.62% β 92.76% | 55.83% β 73.67% | 6.74% β 9.32% | 0.4778 β 0.6373 |
| 10 | 91.14% β 94.27% | 54.47% β 70.49% | 12.83% β 16.99% | 0.4903 β 0.6377 |
| 20 | 95.03% β 95.46% | 52.21% β 64.80% | 23.63% β 29.63% | 0.5201 β 0.6489 |
| 50 | 97.62% β 97.62% | 42.40% β 42.40% | 45.83% β 45.83% | 0.5598 β 0.6125 |
At three results, useful-passage precision changes from 56.04% to 75.11%; nDCG@10 changes from 0.4903 to 0.6377. At 50, Hit, precision and recall are identical because every candidate is returned. nDCG@50 still measures order: 0.5598 β 0.6125. The pool's 97.62% Hit@50 is a retrieval ceiling, not an Ettin accuracy score.
All three variants use matched PyTorch FP16 reranking. Depth 1 uses 919 shared judged queries across the three pipeline variants; depths 3β50 use all 925. Abstentions are excluded inside the original cutoff without backfill. Grades 2β3 count as useful. Precision pools retained positions; recall counts known useful passages. The expanded reference and matched runtime distinguish this table from the earlier CUDA/Vulkan serving check. Daecore returns 3β20 results; these fixed cutoffs do not evaluate its selector.
The evaluation companion provides the paired records, methods and separate serving checks. The pipeline overview explains retrieval and selection.
Training data and objective
Ettin learns to rank prepared document chunks, including useful partial evidence. Several passages can help answer the same question.
| Training ingredient | Purpose |
|---|---|
| Mostly synthetic organizational documents, plus one person's largely AI-written project docs | Supply the notes, procedures and technical material that agents search |
| Frozen structural chunks with section context and table/code handling | Keep supervision tied to the exact passages the model reads |
| Full questions plus 294 answerable lexical-supplement queries | Cover keywords, exact lookups, identifiers, pasted fragments, shorthand and typos |
| Model judgments from 0 to 3, with rotating candidate lists | Compare useful, partly useful and unhelpful passages throughout training |
ListNet learns ordering within each list; a pairwise term emphasizes neighboring grades. The chunking guide explains passage preparation. Earlier parser outputs remain in the frozen corpus, and Ettin applies its tokenizer and pair-length limit afterward.
Training recipe and exposure
| Item | Value |
|---|---|
| Starting model | cross-encoder/ettin-reranker-150m-v1 |
| Training queries | 2,882 with ranking supervision; 30 no-positive queries excluded from the ranking loss |
| Unique queryβpassage pairs | 144,100 |
| Unique passage texts | 35,006 |
| Supervision | Four relevance grades, 0β3; grades 2β3 count as useful |
| Objective | ListNet with an adjacent-grade pairwise term weighted 0.2 |
| Adapter | LoRA+, rank 16, alpha 32; merged for serving |
| Learning rates | 5e-5 for LoRA A and the scoring head; 8e-4 for LoRA B |
| Training steps | 2,882 scheduled; 2,881 applied updates, accumulating four lists of 16 per step |
| Exposure | 11,528 lists; 184,448 candidate presentations |
Runtime and limits
model.onnx is the CUDA FP16 graph and model.vulkan.onnx the Vulkan FP32 graph; both read model.onnx.data. The Vulkan derivation rewrites bounded mask operations and keeps the learned weights. The FP32 graph's results were also checked for correctness on the CPU execution provider; no CPU latency guidance is given.
Vulkan requires onnxruntime==1.24.4 and the native WebGPU plugin onnxruntime-ep-webgpu==0.4.0, registered explicitly. Add the plugin device to the session options with dawnBackendType=Vulkan, and create the session without a providers list; the plugin name alone does not select Vulkan. serving.vulkan.json records the tested provider options.
On the shared 925-query answerable subset, CUDA and Vulkan produced the same Hit@5, Hit@10, Hit@20 and Hit@50; Hit@3 differed by one query and nDCG@10 by 0.0004. The paths are not bit-exact, so close scores can swap order. Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available.
Inside Daecore, Ettin reorders up to 50 deduplicated BM25 and Gemma candidates; the retrieval pipeline overview describes that path and result selection.
The task-specific panel is mostly generated and model-labeled, with no human reference, and was reused during development; training-seed variation is unmeasured. A reranker cannot recover a passage missing from its candidates, and several relevant passages may repeat one fact. These results do not establish unique-evidence coverage or downstream agent success.
License
Apache-2.0. The package includes upstream attribution and a modification notice.