quotebound-27b / evidence_faithful_reasoning_release_brief.md
darcar0's picture
Simplify public release surfaces around the technical note
b314f67 verified
|
Raw History Blame Contribute Delete
4.79 kB
# Quotebound 27B
## Evidence-Faithful Reasoning Release Brief
Released: 2026-04-07
Author: darcar0
Hugging Face model release:
[`darcar0/quotebound-27b`](https://huggingface.co/darcar0/quotebound-27b)
Companion files:
- [`technical_note_evidence_faithful_reasoning.md`](./technical_note_evidence_faithful_reasoning.md)
- [`standalone_holdout_comparison.svg`](./standalone_holdout_comparison.svg)
- [`benchmark_progression.svg`](./benchmark_progression.svg)
## Executive summary
Quotebound 27B is the standalone model release from Evidence-Faithful
Reasoning: a research-engineering project on reasoning that has to stay
recoverable from the source text rather than asserted on top of it. The
project ships a strict benchmark, the hybrid system that clears that
benchmark under the full contract, and a standalone model trained to carry
the same behavior on its own.
The contract is strict. On every closed packet of source text, the system
has to:
1. answer correctly,
2. cite the right evidence units,
3. quote those units verbatim, and
4. abstain with `Insufficient evidence.` when the packet does not justify a
claim.
The project ends in two finished results from one frame. The
benchmark-facing winner is a hybrid stack — bridge `checkpoint-2` plus
`deterministic_v3` packet-local quote normalization — that clears every
gate on the frozen held-out `probe_v0` benchmark. The downloadable artifact
is Quotebound 27B on Hugging Face, which beats the earlier bridge model on
a fresh mixed public holdout and roughly doubles raw quote-faithful
behavior at the model level.
## Quotebound 27B
Quotebound 27B is the strongest standalone model the project produced and
the artifact most readers will load first. It is the first standalone
checkpoint in the project to hold up across multiple evaluation surfaces
beyond the held-out probe.
Fresh 36-task mixed public holdout:
| Stack | Task | Strict | Evidence F1 | Quote F1 |
|---|---:|---:|---:|---:|
| Bridge raw | 0.8611 | 0.2222 | 0.8815 | 0.3343 |
| Quotebound raw | 0.8889 | 0.4444 | 0.9093 | 0.6815 |
| Bridge + `deterministic_v3` | 0.8611 | 0.5833 | 0.8815 | 0.8815 |
| **Quotebound + `deterministic_v3`** | **0.8889** | **0.5833** | **0.9093** | **0.9093** |
Quotebound 27B beats the prior bridge model on task accuracy, evidence F1,
and quote F1 in both raw and normalized form, ties normalized strict, and
roughly doubles raw quote F1 (`0.3343` → `0.6815`) at the model level.
## Benchmark-facing winner
The benchmark-facing hybrid stack is the strongest full system from the
project. Training moved the model past the older frozen baseline; the
deterministic packet-local normalizer was the finishing repair, closing the
remaining quote-faithful gap without leaving the closed-packet boundary.
| Metric | Frozen probe_v0 |
|---|---:|
| Task success | **1.0000** |
| Strict grounded success | **1.0000** |
| Mean evidence F1 | **1.0000** |
| Mean quote F1 | **1.0000** |
| Verify label accuracy | **1.0000** |
| Grounded QA accuracy | **1.0000** |
| Contrastive consistency | **1.0000** |
| Invalid / missing rate | **0.0000** |
## Release boundary
The release has two public faces: Quotebound 27B, the standalone model that
loads directly from Hugging Face, and a benchmark-facing hybrid stack that
closes the last quote-faithfulness gap on the frozen held-out probe. The
split is part of the project story, not hidden behind the fine print.
The public release stops at the point where the strongest benchmark-facing
system and the strongest standalone model were both clearly in hand. That
keeps the package centered on finished artifacts rather than on local
variant history.
## Intended use and boundaries
This is a release for reasoning over closed packets of source text, not a
general-purpose chatbot replacement. It is built for bounded document QA,
claim verification, policy and compliance review, contract reading, and
other settings where every answer has to be justified from a fixed body of
text.
Important boundaries:
- Perfect `probe_v0` belongs to the hybrid stack, not to the standalone
adapter alone.
- The Hugging Face download is the LoRA adapter only; the benchmark-winning
configuration is adapter + `deterministic_v3`.
- Frozen `probe_v0` item-level contents are intentionally not published with
the release.
## Release surfaces
- Quotebound 27B on Hugging Face:
[`darcar0/quotebound-27b`](https://huggingface.co/darcar0/quotebound-27b)
- Technical note:
[`technical_note_evidence_faithful_reasoning.md`](./technical_note_evidence_faithful_reasoning.md)
- Fresh public holdout chart:
[`standalone_holdout_comparison.svg`](./standalone_holdout_comparison.svg)
- Frozen benchmark progression chart:
[`benchmark_progression.svg`](./benchmark_progression.svg)