quotebound-27b / evidence_faithful_reasoning_release_brief.md
darcar0's picture
Simplify public release surfaces around the technical note
b314f67 verified
|
Raw History Blame Contribute Delete
4.79 kB

Quotebound 27B

Evidence-Faithful Reasoning Release Brief

Released: 2026-04-07 Author: darcar0

Hugging Face model release: darcar0/quotebound-27b

Companion files:

Executive summary

Quotebound 27B is the standalone model release from Evidence-Faithful Reasoning: a research-engineering project on reasoning that has to stay recoverable from the source text rather than asserted on top of it. The project ships a strict benchmark, the hybrid system that clears that benchmark under the full contract, and a standalone model trained to carry the same behavior on its own.

The contract is strict. On every closed packet of source text, the system has to:

  1. answer correctly,
  2. cite the right evidence units,
  3. quote those units verbatim, and
  4. abstain with Insufficient evidence. when the packet does not justify a claim.

The project ends in two finished results from one frame. The benchmark-facing winner is a hybrid stack — bridge checkpoint-2 plus deterministic_v3 packet-local quote normalization — that clears every gate on the frozen held-out probe_v0 benchmark. The downloadable artifact is Quotebound 27B on Hugging Face, which beats the earlier bridge model on a fresh mixed public holdout and roughly doubles raw quote-faithful behavior at the model level.

Quotebound 27B

Quotebound 27B is the strongest standalone model the project produced and the artifact most readers will load first. It is the first standalone checkpoint in the project to hold up across multiple evaluation surfaces beyond the held-out probe.

Fresh 36-task mixed public holdout:

Stack Task Strict Evidence F1 Quote F1
Bridge raw 0.8611 0.2222 0.8815 0.3343
Quotebound raw 0.8889 0.4444 0.9093 0.6815
Bridge + deterministic_v3 0.8611 0.5833 0.8815 0.8815
Quotebound + deterministic_v3 0.8889 0.5833 0.9093 0.9093

Quotebound 27B beats the prior bridge model on task accuracy, evidence F1, and quote F1 in both raw and normalized form, ties normalized strict, and roughly doubles raw quote F1 (0.3343 → 0.6815) at the model level.

Benchmark-facing winner

The benchmark-facing hybrid stack is the strongest full system from the project. Training moved the model past the older frozen baseline; the deterministic packet-local normalizer was the finishing repair, closing the remaining quote-faithful gap without leaving the closed-packet boundary.

Metric Frozen probe_v0
Task success 1.0000
Strict grounded success 1.0000
Mean evidence F1 1.0000
Mean quote F1 1.0000
Verify label accuracy 1.0000
Grounded QA accuracy 1.0000
Contrastive consistency 1.0000
Invalid / missing rate 0.0000

Release boundary

The release has two public faces: Quotebound 27B, the standalone model that loads directly from Hugging Face, and a benchmark-facing hybrid stack that closes the last quote-faithfulness gap on the frozen held-out probe. The split is part of the project story, not hidden behind the fine print.

The public release stops at the point where the strongest benchmark-facing system and the strongest standalone model were both clearly in hand. That keeps the package centered on finished artifacts rather than on local variant history.

Intended use and boundaries

This is a release for reasoning over closed packets of source text, not a general-purpose chatbot replacement. It is built for bounded document QA, claim verification, policy and compliance review, contract reading, and other settings where every answer has to be justified from a fixed body of text.

Important boundaries:

  • Perfect probe_v0 belongs to the hybrid stack, not to the standalone adapter alone.
  • The Hugging Face download is the LoRA adapter only; the benchmark-winning configuration is adapter + deterministic_v3.
  • Frozen probe_v0 item-level contents are intentionally not published with the release.

Release surfaces