dispositio v5
This card describes v5, published as the revision
v5(the weights, 1.4 GB). The files onmainare still v3 (the 144M Laya model), kept so an older Inventio that readsmainstill loads what it can read; v4 is at revisionv4, v3's own card atv3. Load v5 withsnapshot_download("minhquan2310/dispositio", revision="v5"), or through Inventio 0.5, which readsminhquan2310/dispositio@v5by default.
Dispositio is the second canon of classical rhetoric: after inventio finds the material, dispositio puts it in order. This model is the local ranker, category judge and type head of Inventio, a retrieval tool for code and documents that uses BM25 and structure instead of embeddings.
v5 is a System One decision model: Kev 0.8B (Qwen3.5-0.8B-Base with a decision head) fine-tuned on Inventio's data, weights merged, 1.4 GB in bf16. It reads one state (a question and up to 15 candidate passages whose lines are numbered) and answers every question asked about it in one pass:
- where: which passage holds the answer, and which line;
- exists: does any passage answer at all;
- category (one passage in the state): Rule, Procedure, Reference, Explanation, Finding, Record or Other;
- type (the question alone): which kind of document would hold the answer.
v5 is v4 trained further on states where the answer is a page another page links to, and on
Vietnamese law. It lives on the tag v5; v4 stays on v4, v3 (the per-passage Laya model) on
v3, and main keeps v3's files so an older Inventio that reads main still loads.
pip install "inventio[dispositio]"
inventio init ~/code/webshop --name app
inventio query "the nightly backup has not finished, what do I do?" # ranked by this model
inventio facts --source app # categories, judged by it
inventio model --use minhquan2310/dispositio@v5 # only on Inventio 0.4; 0.5 reads v5 by default
The state format, the question wordings and the reader are Inventio's (inventio/systemone.py,
inventio/_systemone, Kev's serving path vendored under Apache-2.0); other wordings were not trained.
A state longer than 6,656 tokens was never seen in training.
Results
All numbers are from this checkpoint (s1-v1.4) against v4 (s1-v1.3) on the same questions and
pools, on an RTX 5070 laptop GPU. Intervals are paired per question, bootstrap 95%.
Following links. When a question's answer is on a page the best page links to, Inventio hands the ranker the linked pages beside BM25's best (HotpotQA bridge questions, dev, 300, over their own 66,581 pages, each page's first mention of another page's title drawn as a link). nDCG@10:
| BM25 alone | v4 | v5 | |
|---|---|---|---|
| BM25's 15 only | 0.718 | 0.759 | 0.799 (+0.040 [+0.028, +0.053]) |
| with the linked pages read (the default query) | — | 0.846 | 0.900 (+0.054 [+0.040, +0.069]) |
v4 already gained +0.081 from the links over a same-size BM25 pool, untrained for it; v5 gains more. Training used HotpotQA train questions over a map built from train contexts only.
The passage that answers, ranked first (among questions whose answer BM25 put in the pool):
| Set | Questions | BM25 | v4 | v5 | v5 exists AUC, answer absent (v4) |
|---|---|---|---|---|---|
| MultiDoc2Dial (student aid domain held out whole) | 453 | 0.375 | 0.638 | 0.664 (+0.027 [−0.004, +0.057]) | 0.820 (0.805), 160 pools |
| TechQA (IBM technotes, never trained on) | 87 | 0.149 | 0.322 | 0.345 (+0.023 [−0.058, +0.103]) | 0.875 (0.895), 32 pools |
| examples/webshop (13 on-call questions, written after training) | 13 | 6/13 | 9/13 | 11/13 | — |
- Zalo legal (Vietnamese, 200 test questions, nDCG@10): BM25's 15 only 0.832 against v4 0.818 (+0.014 [−0.005, +0.033]); with names, links and shared-word neighbours read 0.842 against 0.810 (+0.032 [+0.011, +0.055]).
- SciFact, MultiDoc2Dial and TechQA as nDCG@10 (200 queries each, TechQA 119) move by less than 0.02 either way, every interval across zero.
- The right line: on MultiDoc2Dial the top line is inside the answer for 0.552 of the questions (v4 0.561).
- Time to read one pool: median 169 ms on MultiDoc2Dial (3,270 tokens), 379 ms on TechQA (6,480).
Categories, against the labelling model's choice on 380 held-out passages: accuracy / macro-F1 0.811 / 0.774 (v4 0.811 / 0.769; v3 0.663 / 0.642).
Types: not re-measured; v4 put the file a SWE-bench Lite fix changes in the pool for 0.793 of 300 issues.
Training
A continuation of v4: one epoch, 590 optimizer steps over 4,718 records, 57 minutes on one H100
(Modal), LoRA on top of v4's adapter, then merged. recipe.json pins the Kev commit, the base
revision and the data's sha256.
| Source | Records | Label |
|---|---|---|
| HotpotQA train bridge questions over a linked map of their contexts | 1,603 | the two supporting pages; states are BM25's 15 and the final pass Inventio reads, BM25's best beside the pages the top hits link to |
| MultiDoc2Dial train dialogues, three domains (student aid held out) | 1,188 | human: the grounding section |
| Zalo legal train split (test queries removed) | 1,327 | human: the relevant article |
| Category passages, replayed from v4's training | 600 | a small general LLM (Gemini Flash), soft target 0.94 / 0.01 |
- Each state of BM25's 15 has a twin with the answer swapped out, so
existsis also taught "no". - Half the queries have the passages' order shuffled in the state (both twins alike), as RankZephyr does, so the ranking reads content rather than position.
- Only 28% of the Zalo records fit the 6,656-token state: Vietnamese law runs long.
Reproduce with benchmarks/systemone.py data --sources md2d,zalo,hotpot --shuffle 0.5, train
(--bal-source, --judge-sample), and benchmarks/modal_train.py --init-from s1-v1.3.
Limitations
existsis calibrated per corpus: do not read 0.5 as a threshold.- Inventio drops passages from the tail of a pool longer than 26,000 characters before the model reads it (about 6,700 tokens of English, 7,100 of Vietnamese); on Zalo that cost the answer on 23 of 200 test questions with v4's chunking.
- It was not trained to judge whether two passages are about the same thing (Inventio's
aboutlinks) and links almost none; links Inventio shows are drawn from the sources. - Code: no code retrieval states in training (their packed length does not fit an 8 GB card); what it knows of ranking code is Kev's.
- Category labels are one model's reading, not a gold set.
Terms
Apache-2.0, like Kev and Qwen3.5. The training data carries its own terms: MultiDoc2Dial Apache-2.0; HotpotQA CC BY-SA 4.0; SciFact claims CC BY 4.0 and abstracts ODC-By 1.0; Zalo legal (card: MIT); SWE-bench MIT, with each repository's text under its own licence; category labels generated with Gemini.