darcar0 commited on
Commit
9c1d54d
·
verified ·
1 Parent(s): 3f0f458

Upload evidence_faithful_reasoning_release_brief.md with huggingface_hub

Browse files
evidence_faithful_reasoning_release_brief.md ADDED
@@ -0,0 +1,118 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Evidence-Faithful Reasoning
2
+
3
+ ## Release Brief
4
+
5
+ Released: 2026-04-07
6
+ Author: darcar0
7
+
8
+ Hugging Face model release:
9
+ [`darcar0/evidence-faithful-reasoning-pilot-3`](https://huggingface.co/darcar0/evidence-faithful-reasoning-pilot-3)
10
+
11
+ Public GitHub release repo: `PUBLIC_REPO_URL`
12
+
13
+ ## Executive summary
14
+
15
+ I built this project because I wanted a model release where reasoning had to
16
+ cash out into recoverable evidence instead of hiding behind fluent language.
17
+ The result is a strict evidence-faithful reasoning benchmark, a
18
+ benchmark-winning hybrid system, and pilot 3: the strongest standalone model
19
+ from the same project, released as a LoRA adapter on Hugging Face.
20
+
21
+ The contract is strict. On every bounded evidence packet, the system has to:
22
+
23
+ 1. answer correctly,
24
+ 2. identify the right evidence,
25
+ 3. quote the exact supporting text, and
26
+ 4. abstain with `Insufficient evidence.` when the packet does not justify a
27
+ claim.
28
+
29
+ This release ships two finished outputs. The benchmark-facing winner is a
30
+ hybrid stack — bridge `checkpoint-2` plus `deterministic_v3` packet-local
31
+ quote normalization — that clears every gate on the frozen held-out
32
+ `probe_v0` benchmark. The main downloadable artifact is pilot 3, the
33
+ standalone model release, which beats the earlier bridge model on a fresh
34
+ mixed public holdout and roughly doubles raw quote-faithful behavior at the
35
+ model level.
36
+
37
+ ## Benchmark-facing winner
38
+
39
+ The benchmark-facing winner is the strongest full system from the project.
40
+ Training moved the model past the older frozen baseline. Deterministic
41
+ packet-local normalization closed the remaining quote-faithfulness and strict
42
+ grounded-success gap without leaving the bounded-packet setting.
43
+
44
+ | Metric | Frozen probe_v0 |
45
+ |---|---:|
46
+ | Task success | **1.0000** |
47
+ | Strict grounded success | **1.0000** |
48
+ | Mean evidence F1 | **1.0000** |
49
+ | Mean quote F1 | **1.0000** |
50
+ | Verify label accuracy | **1.0000** |
51
+ | Grounded QA accuracy | **1.0000** |
52
+ | Contrastive consistency | **1.0000** |
53
+ | Invalid / missing rate | **0.0000** |
54
+
55
+ Canonical memo:
56
+ [`reports/sft_v1_final_artifact_status.md`](../reports/sft_v1_final_artifact_status.md)
57
+
58
+ ## Standalone model release: pilot 3
59
+
60
+ Pilot 3 is the strongest standalone model from the project and the release I
61
+ want people to download first. It is the first standalone checkpoint in the
62
+ project to hold up across multiple non-`probe_v0` evaluation surfaces, and it
63
+ moves a meaningful amount of the winning behavior into the model itself.
64
+
65
+ Fresh 36-task mixed public holdout:
66
+
67
+ | Stack | Task | Strict | Evidence F1 | Quote F1 |
68
+ |---|---:|---:|---:|---:|
69
+ | Bridge raw | 0.8611 | 0.2222 | 0.8815 | 0.3343 |
70
+ | Pilot 3 raw | 0.8889 | 0.4444 | 0.9093 | 0.6815 |
71
+ | Bridge + `deterministic_v3` | 0.8611 | 0.5833 | 0.8815 | 0.8815 |
72
+ | **Pilot 3 + `deterministic_v3`** | **0.8889** | **0.5833** | **0.9093** | **0.9093** |
73
+
74
+ Pilot 3 beats bridge on task accuracy, evidence F1, and quote F1 in both raw
75
+ and normalized form, ties normalized strict, and roughly doubles raw quote F1
76
+ (`0.3343` → `0.6815`) at the model level.
77
+
78
+ Canonical memos:
79
+ - [`reports/standalone_model_v2_freeze_memo.md`](../reports/standalone_model_v2_freeze_memo.md)
80
+ - [`reports/standalone_model_v2_holdout_v1_bridge_vs_pilot3_status.md`](../reports/standalone_model_v2_holdout_v1_bridge_vs_pilot3_status.md)
81
+
82
+ ## Project arc and stopping point
83
+
84
+ The project followed a full research-engineering loop: baseline mapping,
85
+ benchmark design, prompt and structure interventions, a training-backed bridge
86
+ model, deterministic quote normalization, and then a teacher-student
87
+ distillation cycle to move the winning behavior into the model itself.
88
+
89
+ A targeted follow-up, pilot 4, fixed one specific FEVER
90
+ month/date temporal-insufficiency case but weakened broader behavior on larger
91
+ evaluation surfaces. I treated that as a stop signal rather than churning for
92
+ one more pilot. The release froze at the point where the standalone model was
93
+ strongest and the benchmark-facing system was already complete.
94
+
95
+ ## Intended use and boundaries
96
+
97
+ This is a specialized grounded reasoning release, not a general-purpose
98
+ chatbot replacement. It is built for bounded document QA, claim verification,
99
+ policy/compliance workflows, and other settings where every answer has to be
100
+ justified from a closed body of text.
101
+
102
+ Important boundaries:
103
+
104
+ - Perfect `probe_v0` belongs to the hybrid stack, not to pilot 3 alone.
105
+ - The Hugging Face download is the LoRA adapter only; the benchmark-winning
106
+ configuration is adapter + `deterministic_v3`.
107
+ - Frozen `probe_v0` item-level contents are intentionally not published with
108
+ the release.
109
+
110
+ ## Release surfaces
111
+
112
+ - Pilot 3 on Hugging Face:
113
+ [`darcar0/evidence-faithful-reasoning-pilot-3`](https://huggingface.co/darcar0/evidence-faithful-reasoning-pilot-3)
114
+ - Technical note:
115
+ [`docs/technical_note_evidence_faithful_reasoning.md`](./technical_note_evidence_faithful_reasoning.md)
116
+ - Release page:
117
+ [`release/index.html`](../release/index.html)
118
+ - Public GitHub release repo: `PUBLIC_REPO_URL`