|
Download evaluation/chunking.md from Daecore/ettin-150m-memory-reranker-ft-v1: direct link, hf CLI and curl.
- Browser
- Download file 3.51 kB
-
https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1/resolve/main/evaluation/chunking.md
- Command line
-
hf download hf://Daecore/ettin-150m-memory-reranker-ft-v1/evaluation/chunking.md
-
curl -L -o chunking.md https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1/resolve/main/evaluation/chunking.md
3.51 kB
| # Chunking and retrieval data | |
| ## Retrieval-model data | |
| Daecore fine-tunes retrieval models on passages prepared for its document | |
| workflow. A passage is the exact chunk the model reads: a section, excerpt, | |
| table fragment or code block. It can supply useful partial evidence without | |
| answering the whole question. | |
| **Source documents → structural chunks → graded query–passage pairs → retrieval training** | |
| | Ingredient | What it contributes | | |
| |---|---| | |
| | Mostly synthetic workplace documentation | Notes, procedures, decisions and technical material across generated project scenarios | | |
| | One person's project documents, much of them AI-written | Additional examples, not a representative sample of other users' workspaces | | |
| | Frozen structural chunks | The retrieval units that Gemma embeds and Ettin ranks; some retain earlier parser outputs | | |
| | Model-written questions and graded evidence | Several useful passages per question, with decisive and partial evidence distinguished from merely related text | | |
| | Public evidence annotations used by Gemma v2 | HotpotQA paragraphs, MultiDoc2Dial grounding passages and FinQA text/table evidence | | |
| Public retrieval data also comes in passages; it is not uniformly made of full | |
| documents. The distinction is their source and preparation. Each model applies | |
| its tokenizer and input limits after chunking, so a chunk boundary and a model's | |
| truncation limit are separate constraints. | |
| This is deliberate specialization of already capable upstream models. The | |
| private comparisons measure gains on this workflow, while public benchmarks | |
| show the general-retrieval cost. Daecore accepted that tradeoff for its intended | |
| use; the cards report both sides. Generated questions and shared preparation | |
| and judging processes do not establish gains for every real user or agent. | |
| Training and evaluation use frozen passage snapshots. Current chunking fixes | |
| do not retroactively change those texts or transfer their relevance labels to | |
| different cuts. | |
| ## Chunking methods | |
| - **Heading-based splitting:** use the document's sections; oversized sections | |
| can split further at their subheadings. This follows the author's structure. | |
| - **Title and heading context:** add the document title and relevant heading | |
| path to section chunks so their subject stays clear. Smaller searchable pieces | |
| retain that metadata but do not automatically repeat it inside every piece. | |
| - **Boundary-aware cuts and overlap:** prefer a nearby sentence end, paragraph | |
| break or line break. Ordinary prose windows repeat up to 200 characters across | |
| cuts to carry nearby context; oversized structural continuations have their | |
| own framing rules. | |
| - **Code-block and table handling:** keep blocks together when they fit. Split | |
| oversized tables with repeated column headers and oversized code blocks with | |
| reopened and closed code fences, so continuations remain readable. | |
| - **Record-aware splitting:** preserve row, field or entry context for CSV, | |
| JSON, JSONL and dated entries. Oversized records are divided with identifying | |
| labels rather than losing their remaining content. | |
| - **Paragraph packing for imported memory files:** group whole paragraphs into | |
| bounded pieces so excerpts usually begin at a paragraph. A single oversized | |
| paragraph falls back to structural splitting. | |
| These methods work together: make supported text and record chunks fit first, | |
| then create bounded searchable pieces for sections that are still too large. | |
| Both operations use the same structural cutter. | |