dispositio v5

This card describes v5, published as the revision v5 (the weights, 1.4 GB). The files on main are still v3 (the 144M Laya model), kept so an older Inventio that reads main still loads what it can read; v4 is at revision v4, v3's own card at v3. Load v5 with snapshot_download("minhquan2310/dispositio", revision="v5"), or through Inventio 0.5, which reads minhquan2310/dispositio@v5 by default.

Dispositio is the second canon of classical rhetoric: after inventio finds the material, dispositio puts it in order. This model is the local ranker, category judge and type head of Inventio, a retrieval tool for code and documents that uses BM25 and structure instead of embeddings.

v5 is a System One decision model: Kev 0.8B (Qwen3.5-0.8B-Base with a decision head) fine-tuned on Inventio's data, weights merged, 1.4 GB in bf16. It reads one state (a question and up to 15 candidate passages whose lines are numbered) and answers every question asked about it in one pass:

  • where: which passage holds the answer, and which line;
  • exists: does any passage answer at all;
  • category (one passage in the state): Rule, Procedure, Reference, Explanation, Finding, Record or Other;
  • type (the question alone): which kind of document would hold the answer.

v5 is v4 trained further on states where the answer is a page another page links to, and on Vietnamese law. It lives on the tag v5; v4 stays on v4, v3 (the per-passage Laya model) on v3, and main keeps v3's files so an older Inventio that reads main still loads.

pip install "inventio[dispositio]"
inventio init ~/code/webshop --name app
inventio query "the nightly backup has not finished, what do I do?"   # ranked by this model
inventio facts --source app                                            # categories, judged by it
inventio model --use minhquan2310/dispositio@v5                        # only on Inventio 0.4; 0.5 reads v5 by default

The state format, the question wordings and the reader are Inventio's (inventio/systemone.py, inventio/_systemone, Kev's serving path vendored under Apache-2.0); other wordings were not trained. A state longer than 6,656 tokens was never seen in training.

Results

All numbers are from this checkpoint (s1-v1.4) against v4 (s1-v1.3) on the same questions and pools, on an RTX 5070 laptop GPU. Intervals are paired per question, bootstrap 95%.

Following links. When a question's answer is on a page the best page links to, Inventio hands the ranker the linked pages beside BM25's best (HotpotQA bridge questions, dev, 300, over their own 66,581 pages, each page's first mention of another page's title drawn as a link). nDCG@10:

BM25 alone v4 v5
BM25's 15 only 0.718 0.759 0.799 (+0.040 [+0.028, +0.053])
with the linked pages read (the default query) — 0.846 0.900 (+0.054 [+0.040, +0.069])

v4 already gained +0.081 from the links over a same-size BM25 pool, untrained for it; v5 gains more. Training used HotpotQA train questions over a map built from train contexts only.

The passage that answers, ranked first (among questions whose answer BM25 put in the pool):

Set Questions BM25 v4 v5 v5 exists AUC, answer absent (v4)
MultiDoc2Dial (student aid domain held out whole) 453 0.375 0.638 0.664 (+0.027 [−0.004, +0.057]) 0.820 (0.805), 160 pools
TechQA (IBM technotes, never trained on) 87 0.149 0.322 0.345 (+0.023 [−0.058, +0.103]) 0.875 (0.895), 32 pools
examples/webshop (13 on-call questions, written after training) 13 6/13 9/13 11/13 —
  • Zalo legal (Vietnamese, 200 test questions, nDCG@10): BM25's 15 only 0.832 against v4 0.818 (+0.014 [−0.005, +0.033]); with names, links and shared-word neighbours read 0.842 against 0.810 (+0.032 [+0.011, +0.055]).
  • SciFact, MultiDoc2Dial and TechQA as nDCG@10 (200 queries each, TechQA 119) move by less than 0.02 either way, every interval across zero.
  • The right line: on MultiDoc2Dial the top line is inside the answer for 0.552 of the questions (v4 0.561).
  • Time to read one pool: median 169 ms on MultiDoc2Dial (3,270 tokens), 379 ms on TechQA (6,480).

Categories, against the labelling model's choice on 380 held-out passages: accuracy / macro-F1 0.811 / 0.774 (v4 0.811 / 0.769; v3 0.663 / 0.642).

Types: not re-measured; v4 put the file a SWE-bench Lite fix changes in the pool for 0.793 of 300 issues.

Training

A continuation of v4: one epoch, 590 optimizer steps over 4,718 records, 57 minutes on one H100 (Modal), LoRA on top of v4's adapter, then merged. recipe.json pins the Kev commit, the base revision and the data's sha256.

Source Records Label
HotpotQA train bridge questions over a linked map of their contexts 1,603 the two supporting pages; states are BM25's 15 and the final pass Inventio reads, BM25's best beside the pages the top hits link to
MultiDoc2Dial train dialogues, three domains (student aid held out) 1,188 human: the grounding section
Zalo legal train split (test queries removed) 1,327 human: the relevant article
Category passages, replayed from v4's training 600 a small general LLM (Gemini Flash), soft target 0.94 / 0.01
  • Each state of BM25's 15 has a twin with the answer swapped out, so exists is also taught "no".
  • Half the queries have the passages' order shuffled in the state (both twins alike), as RankZephyr does, so the ranking reads content rather than position.
  • Only 28% of the Zalo records fit the 6,656-token state: Vietnamese law runs long.

Reproduce with benchmarks/systemone.py data --sources md2d,zalo,hotpot --shuffle 0.5, train (--bal-source, --judge-sample), and benchmarks/modal_train.py --init-from s1-v1.3.

Limitations

  • exists is calibrated per corpus: do not read 0.5 as a threshold.
  • Inventio drops passages from the tail of a pool longer than 26,000 characters before the model reads it (about 6,700 tokens of English, 7,100 of Vietnamese); on Zalo that cost the answer on 23 of 200 test questions with v4's chunking.
  • It was not trained to judge whether two passages are about the same thing (Inventio's about links) and links almost none; links Inventio shows are drawn from the sources.
  • Code: no code retrieval states in training (their packed length does not fit an 8 GB card); what it knows of ranking code is Kev's.
  • Category labels are one model's reading, not a gold set.

Terms

Apache-2.0, like Kev and Qwen3.5. The training data carries its own terms: MultiDoc2Dial Apache-2.0; HotpotQA CC BY-SA 4.0; SciFact claims CC BY 4.0 and abstracts ODC-By 1.0; Zalo legal (card: MIT); SWE-bench MIT, with each repository's text under its own licence; category labels generated with Gemini.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.1B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for minhquan2310/dispositio

Finetuned
(1)
this model

Datasets used to train minhquan2310/dispositio