tnh0527's picture
Update ettin-150m-memory-reranker-ft-v1 documentation
de87c88 verified
|
Raw History Blame Contribute Delete
3.51 kB
# Chunking and retrieval data
## Retrieval-model data
Daecore fine-tunes retrieval models on passages prepared for its document
workflow. A passage is the exact chunk the model reads: a section, excerpt,
table fragment or code block. It can supply useful partial evidence without
answering the whole question.
**Source documents → structural chunks → graded query–passage pairs → retrieval training**
| Ingredient | What it contributes |
|---|---|
| Mostly synthetic workplace documentation | Notes, procedures, decisions and technical material across generated project scenarios |
| One person's project documents, much of them AI-written | Additional examples, not a representative sample of other users' workspaces |
| Frozen structural chunks | The retrieval units that Gemma embeds and Ettin ranks; some retain earlier parser outputs |
| Model-written questions and graded evidence | Several useful passages per question, with decisive and partial evidence distinguished from merely related text |
| Public evidence annotations used by Gemma v2 | HotpotQA paragraphs, MultiDoc2Dial grounding passages and FinQA text/table evidence |
Public retrieval data also comes in passages; it is not uniformly made of full
documents. The distinction is their source and preparation. Each model applies
its tokenizer and input limits after chunking, so a chunk boundary and a model's
truncation limit are separate constraints.
This is deliberate specialization of already capable upstream models. The
private comparisons measure gains on this workflow, while public benchmarks
show the general-retrieval cost. Daecore accepted that tradeoff for its intended
use; the cards report both sides. Generated questions and shared preparation
and judging processes do not establish gains for every real user or agent.
Training and evaluation use frozen passage snapshots. Current chunking fixes
do not retroactively change those texts or transfer their relevance labels to
different cuts.
## Chunking methods
- **Heading-based splitting:** use the document's sections; oversized sections
can split further at their subheadings. This follows the author's structure.
- **Title and heading context:** add the document title and relevant heading
path to section chunks so their subject stays clear. Smaller searchable pieces
retain that metadata but do not automatically repeat it inside every piece.
- **Boundary-aware cuts and overlap:** prefer a nearby sentence end, paragraph
break or line break. Ordinary prose windows repeat up to 200 characters across
cuts to carry nearby context; oversized structural continuations have their
own framing rules.
- **Code-block and table handling:** keep blocks together when they fit. Split
oversized tables with repeated column headers and oversized code blocks with
reopened and closed code fences, so continuations remain readable.
- **Record-aware splitting:** preserve row, field or entry context for CSV,
JSON, JSONL and dated entries. Oversized records are divided with identifying
labels rather than losing their remaining content.
- **Paragraph packing for imported memory files:** group whole paragraphs into
bounded pieces so excerpts usually begin at a paragraph. A single oversized
paragraph falls back to structural splitting.
These methods work together: make supported text and record chunks fit first,
then create bounded searchable pieces for sections that are still too large.
Both operations use the same structural cutter.