README / README.md
tnh0527's picture
Show each retrieval model’s contribution to the pipeline
c471997 verified
|
Raw History Blame Contribute Delete
2.86 kB
metadata
title: Daecore
sdk: static
emoji: 📚
colorFrom: gray
colorTo: blue

Small models for local search and passage labeling

Daecore's models find and organize evidence in working notes and technical documentation. They run locally through ONNX Runtime, without a hosted model API, and each package can be used on its own.

Model Use it to Returns License
EmbeddingGemma 300M v2 Find passages related to a query A normalized 768-dimensional vector per query or passage Gemma Terms of Use
Ettin 150M Put the most useful candidates first A relevance score per query–passage pair Apache-2.0
GLiClass Knowledge Classifier v2 Label what a passage contains Scores for trap, decision, constraint, mechanism and procedure Apache-2.0

Retrieval. Gemma and Ettin form Daecore's retrieval pipeline. Gemma and BM25 keyword search retrieve candidates, reciprocal-rank fusion merges their rankings, and Ettin reranks up to 50 unique passages. A calibrated selector then returns 3–20 of them when enough candidates exist.

Passages are prepared before retrieval: deterministic structural chunking follows headings and natural boundaries, carries section context, and keeps code and table framing where possible. Token-aware splitting bounds oversized material. Training uses frozen snapshots of these passages, including earlier parser outputs; changing the chunker changes the evidence available to search. Results are source passages for an agent to read and use, and several relevant passages can repeat the same fact.

Passage labeling. GLiClass serves a separate Daecore function: labeling what a passage contains, so stored material can be organized by kind. It is not a retrieval component and needs neither retrieval model; its labels do not judge relevance, truth or authority.

Each card has a runnable example near the top, its training recipe and a comparison with the exact upstream checkpoint on Daecore's evaluation data. Gemma also reports five public retrieval datasets and the storage/quality tradeoff of smaller embeddings; Ettin reports two public reranking datasets. Their pipeline tables change one model at a time against the same fully fine-tuned system, showing what each contributes alongside BM25.

Task-specific training and evaluation rely on mostly generated organizational documents and model-generated labels, without a human-adjudicated reference, so reported gains describe that data rather than a general improvement over the upstream models. Serving checks ran on one Windows x64 machine with an NVIDIA RTX 3060 Ti; Linux and AMD or Intel GPUs are untested.