# Technical Note: Evidence-Faithful Reasoning Over Bounded Evidence Packets Updated: 2026-04-07 ## Abstract This note describes a research-engineering project on strict evidence-faithful reasoning in open reasoning-distilled language models. The system answers from a closed packet of source text and must satisfy four conditions at once: answer correctly, identify the right evidence units, quote exact supporting text, and abstain with `Insufficient evidence.` when the packet does not justify a claim. The project produced two finished outputs from one frame: a benchmark-winning hybrid system that solves the frozen held-out `probe_v0` benchmark under the full contract, and pilot 3, the strongest standalone model artifact from the project, released as a LoRA adapter on Hugging Face. A narrowly targeted follow-up checkpoint, pilot 4, was rejected as a stop signal after it traded one local fix for broader regressions. The hybrid stack is the benchmark-facing release; pilot 3 is the standalone model release. **Keywords:** evidence-faithful reasoning, grounded QA, claim verification, bounded evidence packets, abstention, attribution. ## 1. Problem Many language models can produce correct-looking answers while grounding poorly. In practice, that failure shows up as one or more of the following: - the answer is right but the cited evidence is wrong - the quote is too broad, too narrow, or not recoverable from the source text - the model fails to abstain when the packet does not justify the answer - the output sounds persuasive even when the evidence is insufficient The question is not whether the model can answer from a packet. The question is whether it can answer **faithfully** — with recoverable support and the correct abstention behavior when the packet falls short. ## 2. Benchmark contract For each task on a bounded evidence packet, the system must clear all four gates at once: 1. **Answer correctly** — the right answer or label for the task. 2. **Pick the right evidence** — the cited evidence units must be the packet locations that actually support the answer. 3. **Quote exact support** — every quote has to be a verbatim substring of the cited unit; no paraphrase, no stitching, no ellipsis. 4. **Abstain when blocked** — if the packet does not justify a claim, the answer must be exactly `Insufficient evidence.` Correctness alone does not count as success. The frozen held-out benchmark surface for the final release cycle is `data/probe_v0/`. It remained frozen throughout the cycle and was not used as a tuning surface. ## 3. Method The project followed a full research-engineering loop: 1. build the evaluation harness and frozen baselines 2. produce a baseline table and failure mapping 3. test prompt and structure variants 4. build a training-backed intervention path 5. add deterministic packet-local quote normalization where the model still underperformed the strict contract 6. rerun held-out evaluation on a protected split 7. run a teacher-student distillation cycle to push the winning behavior into the model itself 8. stop when the strongest release artifact emerged Train and dev surfaces are public-data-backed and derived from FEVER-style verify-claim data, HotpotQA-style grounded QA data, and project-local packet scaffolding built on top of those upstream sources. ## 4. Benchmark-facing system The benchmark-facing release is a hybrid system: - **Bridge checkpoint:** `outputs/sft_v1_v2_qlora_bridge_v1_run1/checkpoint-2` - **Deterministic packet-local normalization:** `python3 scripts/normalize_quotes.py --mode deterministic_v3` Both parts are necessary. The training step closes the gap between the older frozen baseline and the strict contract on task and evidence accuracy. The deterministic normalizer is the finishing move that closes the contract on quote faithfulness and strict grounded success without leaving the closed packet. ## 5. Benchmark-facing result Frozen held-out `probe_v0` progression: | Stack | Task | Strict | Evidence F1 | Quote F1 | |---|---:|---:|---:|---:| | Frozen `v2` baseline | 0.9545 | 0.1818 | 0.8758 | 0.2494 | | Bridge `checkpoint-2` (raw) | 1.0000 | 0.2727 | 0.8844 | 0.4409 | | Bridge + `deterministic_v2` | 1.0000 | 0.4091 | 0.8844 | 0.5773 | | **Bridge + `deterministic_v3`** | **1.0000** | **1.0000** | **1.0000** | **1.0000** | See [Figure 1: benchmark progression](../release/assets/benchmark_progression.svg). The full final-artifact frozen `probe_v0` score is task `1.0000`, strict grounded success `1.0000`, evidence F1 `1.0000`, quote F1 `1.0000`, verify label accuracy `1.0000`, grounded QA accuracy `1.0000`, contrastive consistency `1.0000`, invalid / missing rate `0.0000`. Canonical memo: `reports/sft_v1_final_artifact_status.md`. ## 6. Standalone model: pilot 3 After the hybrid stack solved the benchmark, the project asked a second question inside the same frame: how much of that winning behavior can be moved into the model itself, evaluated outside `probe_v0`? That question produced pilot 3: a teacher-student distillation checkpoint released as a LoRA adapter on top of [`Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2`](https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2). - **Internal checkpoint:** `outputs/sft_v1_v2_teacher_distill_pilot_v3_partialdev/checkpoint-16` - **Public release identity:** [`darcar0/evidence-faithful-reasoning-pilot-3`](https://huggingface.co/darcar0/evidence-faithful-reasoning-pilot-3) Pilot 3 is the strongest standalone model artifact from the project. It is the first standalone checkpoint to hold up across multiple non-`probe_v0` evaluation surfaces, and it roughly doubles raw quote-faithful behavior over the earlier bridge model on a fresh public holdout. Selection happened entirely off `probe_v0`; the held-out gate stayed frozen. ## 7. Standalone model: results ### 7.1 Fresh 36-task mixed public holdout A held-out slice of 18 FEVER verify-claim tasks plus 18 HotpotQA grounded QA tasks, drawn from public sources and de-duplicated against every training, dev, and `probe_v0` row. | Stack | Task | Strict | Evidence F1 | Quote F1 | |---|---:|---:|---:|---:| | Bridge raw | 0.8611 | 0.2222 | 0.8815 | 0.3343 | | Pilot 3 raw | 0.8889 | 0.4444 | 0.9093 | 0.6815 | | Bridge + `deterministic_v3` | 0.8611 | 0.5833 | 0.8815 | 0.8815 | | **Pilot 3 + `deterministic_v3`** | **0.8889** | **0.5833** | **0.9093** | **0.9093** | See [Figure 2: standalone holdout comparison](../release/assets/standalone_holdout_comparison.svg). Pilot 3 beats bridge on task accuracy, evidence F1, and quote F1 in both raw and normalized form, ties normalized strict, and roughly doubles raw quote F1 at the model level. Grounded-QA accuracy on this slice is `1.0000` for both stacks; zero invalid outputs across every reported evaluation surface. Canonical memo: `reports/standalone_model_v2_holdout_v1_bridge_vs_pilot3_status.md`. ### 7.2 Fixed dev triage slice (21 tasks) | Stack | Task | Strict | Evidence F1 | Quote F1 | |---|---:|---:|---:|---:| | Pilot 3 + `deterministic_v3` | 1.0000 | 0.6190 | 0.8320 | 0.7095 | ### 7.3 Untouched 104-task Hotpot shadow slice Pilot 3 raw improved quote-faithful behavior over raw bridge, and pilot 3 + `deterministic_v3` matched bridge + `deterministic_v3` at the system level on this slice. Reported as a parity outcome in the standalone freeze memo (`reports/standalone_model_v2_freeze_memo.md`); per-metric numbers were not recorded for this surface, so it stands as a narrative parity result, not a table cell. ## 8. Stop signal: pilot 4 A targeted follow-up, pilot 4, was built to fix one specific FEVER month/date temporal-insufficiency error. It fixed that single row but weakened broader behavior on the larger evaluation surfaces. That outcome made pilot 4 useful as a stop signal: further local-fix iteration was trading visible gains for wider regressions. The project froze at pilot 3 as the strongest standalone model. Canonical memo: `reports/sft_v1_v2_teacher_distill_pilot_v4_partialdev_status.md`. See [Figure 3: project release arc](../release/assets/project_release_arc.svg) for the full arc from baseline through pilot 4 stop signal. ## 9. Discussion **Why the release shape is hybrid + standalone, not one or the other.** The strict contract has four gates. Training alone closed the task and evidence gates on `probe_v0`, but bridge `checkpoint-2` did not close the strict grounded success gate by itself: raw bridge was at strict `0.2727` on `probe_v0`. Deterministic packet-local normalization was the finishing move that closed the remaining gates without escaping the closed-packet boundary. That is what made the hybrid stack the strongest benchmark-facing result and why it is the benchmark-facing release. **What pilot 3 demonstrates.** The teacher-student cycle that produced pilot 3 asked how much of that winning behavior could move into the model itself, evaluated entirely off `probe_v0`. The fresh-holdout deltas show that the model-side gain is real: raw pilot 3 roughly doubles raw bridge on quote F1 (`0.3343` → `0.6815`) and beats raw bridge on task accuracy and evidence F1. This is the first standalone model in the project that holds up across multiple non-`probe_v0` surfaces; calling it the standalone release artifact is faithful to that evidence. **Why pilot 4 is informative.** Pilot 4 was a deliberate narrow refinement. It worked on the one row it targeted and regressed broader behavior. Reading that result as a stop signal — rather than running additional local fixes — is the discipline the project decided to keep visible in the release trail. **Distinction held throughout.** Perfect frozen `probe_v0` belongs to the hybrid stack, not pilot 3 alone. The release ships both results without collapsing them into one claim. ## 10. Contributions 1. A strict packet-faithful reasoning benchmark setup that requires answer, evidence, exact quotes, and `Insufficient evidence.` abstention together, on a frozen held-out surface. 2. A baseline-to-intervention evidence trail that documents what training, prompt structure, and deterministic packet-local normalization each contributed. 3. A finished benchmark-winning hybrid artifact: bridge `checkpoint-2` plus `deterministic_v3` packet-local normalization. 4. A standalone model release — pilot 3, a LoRA adapter on a 27B reasoning-distilled base — that is the first standalone checkpoint to hold up across multiple non-`probe_v0` evaluation surfaces and that roughly doubles raw quote-faithful behavior over the earlier bridge. 5. A stop signal — pilot 4 — that documents the point at which further local-fix iteration began trading visible gains for wider regressions. ## 11. Intended use and limitations **Intended use.** Specialized grounded reasoning over bounded evidence packets: bounded document QA, claim verification, policy and compliance review, contract reading, and other workflows where every answer has to be justified from a closed body of text. Also: research on evidence-faithful reasoning and abstention behavior. **Limitations.** - This is **not** a general-purpose chatbot replacement. Performance outside the closed-packet setting is not characterized. - Perfect `probe_v0` belongs to the **hybrid stack**, not to pilot 3 alone. Treat `1.0000` numbers on `probe_v0` as the hybrid stack's result. - The downloadable artifact for pilot 3 is the LoRA adapter only. `deterministic_v3` is a separate post-processing step that lives in the project repository; the benchmark-winning configuration is *adapter + normalizer*. - Frozen `probe_v0` item-level contents are intentionally not published with the release in order to preserve the held-out gate. - Perfect `probe_v0` is not proof of general faithful reasoning. It is proof that the system meets the strict contract on a single frozen bounded benchmark. ## 12. Surfaces Canonical project surfaces: - final artifact memo: `reports/sft_v1_final_artifact_status.md` - standalone freeze memo: `reports/standalone_model_v2_freeze_memo.md` - fresh holdout comparison: `reports/standalone_model_v2_holdout_v1_bridge_vs_pilot3_status.md` - pilot 4 stop-signal memo: `reports/sft_v1_v2_teacher_distill_pilot_v4_partialdev_status.md` - release brief: `docs/evidence_faithful_reasoning_release_brief.pdf` - release page: `release/index.html` - model card: `release/huggingface_model_card_README.md` - Hugging Face release: [`darcar0/evidence-faithful-reasoning-pilot-3`](https://huggingface.co/darcar0/evidence-faithful-reasoning-pilot-3)