tnh0527's picture
Update ettin-150m-memory-reranker-ft-v1 documentation
de87c88 verified
|
Raw History Blame Contribute Delete
3.51 kB

Chunking and retrieval data

Retrieval-model data

Daecore fine-tunes retrieval models on passages prepared for its document workflow. A passage is the exact chunk the model reads: a section, excerpt, table fragment or code block. It can supply useful partial evidence without answering the whole question.

Source documents → structural chunks → graded query–passage pairs → retrieval training

Ingredient What it contributes
Mostly synthetic workplace documentation Notes, procedures, decisions and technical material across generated project scenarios
One person's project documents, much of them AI-written Additional examples, not a representative sample of other users' workspaces
Frozen structural chunks The retrieval units that Gemma embeds and Ettin ranks; some retain earlier parser outputs
Model-written questions and graded evidence Several useful passages per question, with decisive and partial evidence distinguished from merely related text
Public evidence annotations used by Gemma v2 HotpotQA paragraphs, MultiDoc2Dial grounding passages and FinQA text/table evidence

Public retrieval data also comes in passages; it is not uniformly made of full documents. The distinction is their source and preparation. Each model applies its tokenizer and input limits after chunking, so a chunk boundary and a model's truncation limit are separate constraints.

This is deliberate specialization of already capable upstream models. The private comparisons measure gains on this workflow, while public benchmarks show the general-retrieval cost. Daecore accepted that tradeoff for its intended use; the cards report both sides. Generated questions and shared preparation and judging processes do not establish gains for every real user or agent.

Training and evaluation use frozen passage snapshots. Current chunking fixes do not retroactively change those texts or transfer their relevance labels to different cuts.

Chunking methods

  • Heading-based splitting: use the document's sections; oversized sections can split further at their subheadings. This follows the author's structure.
  • Title and heading context: add the document title and relevant heading path to section chunks so their subject stays clear. Smaller searchable pieces retain that metadata but do not automatically repeat it inside every piece.
  • Boundary-aware cuts and overlap: prefer a nearby sentence end, paragraph break or line break. Ordinary prose windows repeat up to 200 characters across cuts to carry nearby context; oversized structural continuations have their own framing rules.
  • Code-block and table handling: keep blocks together when they fit. Split oversized tables with repeated column headers and oversized code blocks with reopened and closed code fences, so continuations remain readable.
  • Record-aware splitting: preserve row, field or entry context for CSV, JSON, JSONL and dated entries. Oversized records are divided with identifying labels rather than losing their remaining content.
  • Paragraph packing for imported memory files: group whole paragraphs into bounded pieces so excerpts usually begin at a paragraph. A single oversized paragraph falls back to structural splitting.

These methods work together: make supported text and record chunks fit first, then create bounded searchable pieces for sections that are still too large. Both operations use the same structural cutter.