darcar0 commited on
Commit
e5e90d1
·
verified ·
1 Parent(s): 23dce5f

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +92 -142
README.md CHANGED
@@ -21,56 +21,53 @@ tags:
21
  - research
22
  ---
23
 
24
- # Evidence-Faithful Reasoning — Pilot 3
25
 
26
- **Pilot 3 is a LoRA adapter that makes its reasoning-distilled 27B base model
27
- earn every answer from a bounded evidence packet.**
28
 
29
- I built this project because I wanted a release where reasoning had to cash
30
- out into recoverable evidence, not just fluent confidence. In this project,
31
- the model has to point to the right packet units, quote exact support, and
32
- abstain cleanly when the text runs out.
33
-
34
- Given a closed packet of source text and a task, the model has to do four
35
- things at once:
36
-
37
- 1. answer correctly,
38
- 2. identify the right evidence units,
39
- 3. quote the exact supporting text, and
40
- 4. abstain with `Insufficient evidence.` when the packet does not justify
41
- a claim.
42
-
43
- Pilot 3 is the strongest standalone model from that project. The same project
44
- also produced a benchmark-winning hybrid system (bridge `checkpoint-2` plus a
45
- deterministic packet-local normalizer); **this page is the front door for the
46
- standalone model.**
47
 
48
  ## Resources & Guides
49
 
50
  - [Technical brief (PDF)](./evidence_faithful_reasoning_release_brief.pdf)
51
  - [Technical note](./technical_note_evidence_faithful_reasoning.md)
52
  - [Fresh public holdout chart](./standalone_holdout_comparison.svg)
53
- - Public GitHub release repo: `PUBLIC_REPO_URL`
 
 
 
54
 
55
- ![Pilot 3 release proof](./standalone_holdout_comparison.svg)
 
 
56
 
57
- *Fresh 36-task mixed public holdout: pilot 3 beats bridge on task accuracy,
58
- evidence F1, and quote F1, while tying normalized strict after
59
- `deterministic_v3`.*
60
 
61
- ## Release highlights
 
 
 
 
 
 
 
62
 
63
- - **Strongest standalone model from the project.** First standalone
64
- checkpoint to hold up across multiple non-`probe_v0` evaluation
65
- surfaces.
66
- - **Roughly doubles raw quote-faithfulness** over the earlier bridge
67
- model on a fresh public holdout (raw quote F1 `0.3343` → `0.6815`).
68
- - **Beats bridge on a fresh 36-task mixed public holdout** in task
69
- accuracy, evidence F1, and quote F1, while tying normalized strict
70
- grounded success.
71
- - **Zero invalid outputs** on every reported evaluation surface.
72
  - LoRA adapter on top of
73
  [`Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2`](https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2).
 
 
 
 
 
 
 
 
74
 
75
  ## Quick start
76
 
@@ -88,14 +85,12 @@ base = AutoModelForCausalLM.from_pretrained(base_id, device_map="auto")
88
  model = PeftModel.from_pretrained(base, adapter_id)
89
  ```
90
 
91
- The base model is 27B parameters — load it in your usual quantization.
92
 
93
- ### How to prompt it
94
 
95
- Pilot 3 expects an evidence-first strict-contract prompt. The system
96
- message tells the model to find the smallest sufficient evidence set,
97
- copy verbatim quotes from those units, and emit a structured JSON
98
- response. A short version of the recommended prompt:
99
 
100
  ```
101
  You are answering from a bounded evidence packet only.
@@ -129,19 +124,13 @@ The model then writes a JSON object with this shape:
129
  }
130
  ```
131
 
132
- The prompt skeleton above is enough to run the adapter directly. In the
133
- benchmark-winning configuration, the JSON output is then passed through a
134
- deterministic packet-local quote normalizer; *Project context* explains what
135
- that adds and what it does not.
136
-
137
- ## Evaluation results
138
 
139
  ### Fresh 36-task mixed public holdout
140
 
141
- A held-out slice of 18 FEVER verify-claim tasks plus 18 HotpotQA grounded
142
- QA tasks, drawn from public sources and de-duplicated against every
143
- training, dev, and `probe_v0` row. Source: the standalone holdout
144
- comparison report.
145
 
146
  | Stack | Task | Strict | Evidence F1 | Quote F1 |
147
  |---|---:|---:|---:|---:|
@@ -150,10 +139,9 @@ comparison report.
150
  | Bridge + `deterministic_v3` | 0.8611 | 0.5833 | 0.8815 | 0.8815 |
151
  | **Pilot 3 + `deterministic_v3`** | **0.8889** | **0.5833** | **0.9093** | **0.9093** |
152
 
153
- Pilot 3 beats bridge on task accuracy, evidence F1, and quote F1 in both
154
- raw and normalized form, ties normalized strict, and roughly **doubles**
155
- raw quote F1 at the model level. Grounded-QA accuracy on this slice is
156
- `1.0000`.
157
 
158
  ### Fixed dev triage slice (21 tasks)
159
 
@@ -163,110 +151,72 @@ raw quote F1 at the model level. Grounded-QA accuracy on this slice is
163
 
164
  ### Untouched 104-task Hotpot shadow slice
165
 
166
- Pilot 3 raw improved quote-faithful behavior over the raw bridge model on
167
- this slice, and pilot 3 + `deterministic_v3` matched bridge +
168
- `deterministic_v3` at the system level. Reported as a parity outcome in
169
- the standalone freeze memo.
170
 
171
- ## Project context
172
 
173
- The release ships two distinct outputs from one project:
174
 
175
- 1. **Benchmark-winning hybrid system** — bridge `checkpoint-2` plus
176
- `deterministic_v3` packet-local normalization. Perfect on the frozen
177
- held-out `probe_v0` benchmark under the full
178
- answer–evidence–quote–abstain contract: task `1.0000`, strict
179
- grounded success `1.0000`, evidence F1 `1.0000`, quote F1 `1.0000`,
180
- verify accuracy `1.0000`, invalid rate `0.0000`. This is the
181
- strongest full system result from the project.
182
- 2. **Pilot 3 standalone model** — this Hugging Face release. The
183
- strongest version of the project's evidence-faithful behavior that
184
- moved into the model itself, evaluated across multiple non-`probe_v0`
185
- surfaces.
186
 
187
- The two releases are intentionally separate. The hybrid stack is the
188
- benchmark-facing winner. Pilot 3 is the main downloadable model release.
189
- Perfect `probe_v0` belongs to the hybrid stack, not to pilot 3 alone.
 
190
 
191
  ## Intended use
192
 
193
- This is a **specialized grounded reasoning model**, not a general-purpose
194
- chatbot replacement. It is built for:
195
 
196
  - bounded document QA with explicit evidence requirements,
197
- - claim verification and grounded QA from fixed evidence packets,
198
- - policy, compliance, contract, and internal-document workflows where
199
- every answer has to be justified from a closed body of text,
200
  - research on evidence-faithful reasoning and abstention behavior.
201
 
202
- The contract is strict by design: correctness alone does not count as
203
- success. Every answer must come with recoverable support, and the model
204
- must abstain when the packet does not justify a claim.
205
-
206
- ## Training data
207
-
208
- Training and evaluation surfaces are public-data-backed and derived from:
209
-
210
- - **FEVER** — verify-claim data with support / contradict / insufficient
211
- labels.
212
- - **HotpotQA** — multi-hop grounded QA over short evidence packets.
213
- - Project-local bounded packet scaffolding built on top of those
214
- upstream sources.
215
-
216
- The held-out benchmark `probe_v0` was kept frozen and was **not** used as
217
- a tuning surface for the standalone selection cycle that produced this
218
- adapter.
219
-
220
  ## Limitations
221
 
222
- - **The downloadable artifact is the LoRA adapter only.** The
223
- `deterministic_v3` packet-local normalizer is a separate post-processing
224
- step that lives in the project repository. The benchmark-winning
225
- configuration is *adapter + normalizer*; downloading the adapter alone
226
- reproduces the standalone-model numbers, not the perfect hybrid
227
- numbers.
228
- - **Perfect `probe_v0` belongs to the hybrid stack, not to pilot 3
229
- alone.** Treat `1.0000` numbers on `probe_v0` as the hybrid stack's
230
- result.
231
- - **Specialized for bounded evidence packets.** This is not a
232
- general-purpose chatbot or open-domain QA model. Performance outside
233
- the closed-packet setting is not characterized.
234
- - **Frozen benchmark contents are not published with the release.** Raw
235
- `probe_v0` item-level contents are intentionally withheld to preserve
236
- the held-out gate.
237
- - **Perfect `probe_v0` is not proof of general faithful reasoning.** It
238
- is proof that the system meets the strict contract on a single frozen
239
- bounded benchmark.
240
-
241
- ## Why pilot 3 is the release checkpoint
242
-
243
- The project deliberately stopped at pilot 3.
244
-
245
- A targeted follow-up, pilot 4, was built to fix one specific FEVER
246
- month/date temporal-insufficiency error. It fixed that single row, but
247
- it weakened broader behavior on the larger evaluation surfaces. That
248
- made pilot 4 a stop signal — useful negative evidence that further
249
- local-fix iteration was trading visible gains for wider regressions —
250
- not a better release. The project froze at the last point where the
251
- standalone model was strongest across multiple surfaces.
252
-
253
- ## References
254
-
255
- - Base model: [Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2](https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2)
256
- - FEVER: [fever/fever](https://huggingface.co/datasets/fever/fever)
257
- - HotpotQA: [hotpotqa/hotpot_qa](https://huggingface.co/datasets/hotpotqa/hotpot_qa)
258
- - Technical brief (PDF): [evidence_faithful_reasoning_release_brief.pdf](./evidence_faithful_reasoning_release_brief.pdf)
259
- - Technical note: [technical_note_evidence_faithful_reasoning.md](./technical_note_evidence_faithful_reasoning.md)
260
 
261
  ## Citation
262
 
 
 
 
 
 
 
 
 
 
 
 
 
263
  ```bibtex
264
- @misc{evidence_faithful_reasoning_pilot_3_2026,
265
- title = {Evidence-Faithful Reasoning --- Pilot 3},
266
- author = {{darcar0}},
267
  year = {2026},
268
  howpublished = {Hugging Face model release},
269
- url = {https://huggingface.co/darcar0/evidence-faithful-reasoning-pilot-3},
270
- note = {Standalone LoRA adapter from the evidence-faithful reasoning project.}
271
  }
272
  ```
 
21
  - research
22
  ---
23
 
24
+ # Evidence-Faithful Reasoning
25
 
26
+ *Pilot 3 is the standalone model release.*
 
27
 
28
+ This adapter turns its reasoning-distilled 27B base model into an
29
+ evidence-first reader for closed packets of text. I built it because I wanted
30
+ a release where reasoning had to prove itself: every answer has to land on the
31
+ right evidence, quote that evidence verbatim, and stop with
32
+ `Insufficient evidence.` when the packet does not justify a claim. The result
33
+ is the strongest standalone model from the project, packaged here as a LoRA
34
+ adapter you can load directly tonight.
 
 
 
 
 
 
 
 
 
 
 
35
 
36
  ## Resources & Guides
37
 
38
  - [Technical brief (PDF)](./evidence_faithful_reasoning_release_brief.pdf)
39
  - [Technical note](./technical_note_evidence_faithful_reasoning.md)
40
  - [Fresh public holdout chart](./standalone_holdout_comparison.svg)
41
+ - [Frozen benchmark progression chart](./benchmark_progression.svg)
42
+ - [Release architecture chart](./project_release_arc.svg)
43
+
44
+ ![Fresh public holdout: standalone release vs bridge](./standalone_holdout_comparison.svg)
45
 
46
+ *Fresh 36-task mixed public holdout: the standalone release beats the earlier
47
+ bridge model on task accuracy, evidence F1, and quote F1, while the
48
+ packet-local normalizer lifts the full stack to `0.9093` quote F1.*
49
 
50
+ ## Why this release exists
 
 
51
 
52
+ I built this project to force reasoning models to show their work in the only
53
+ place that counts: the evidence itself. Fluent answers were not enough. I
54
+ wanted a model that had to retrieve the right units, quote them exactly, and
55
+ fail closed when the packet ran out. This page leads with the standalone
56
+ release because it is the artifact you can load immediately, inspect directly,
57
+ and use without reconstructing the whole benchmark stack.
58
+
59
+ ## At a glance
60
 
 
 
 
 
 
 
 
 
 
61
  - LoRA adapter on top of
62
  [`Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2`](https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2).
63
+ - Strongest standalone model from the project and the release I want people to
64
+ download first.
65
+ - On a fresh 36-task public holdout, raw task improves from `0.8611` to
66
+ `0.8889`, raw strict from `0.2222` to `0.4444`, and raw quote F1 from
67
+ `0.3343` to `0.6815` over the earlier bridge model.
68
+ - Zero invalid outputs on every reported evaluation surface.
69
+ - The project also produced a benchmark-winning hybrid stack, but that is a
70
+ separate result described under *Release architecture*.
71
 
72
  ## Quick start
73
 
 
85
  model = PeftModel.from_pretrained(base, adapter_id)
86
  ```
87
 
88
+ The base model is 27B parameters, so load it in your usual quantization.
89
 
90
+ ## Prompt format
91
 
92
+ This release works best with an evidence-first prompt that makes the answer
93
+ subordinate to the cited text. A minimal version:
 
 
94
 
95
  ```
96
  You are answering from a bounded evidence packet only.
 
124
  }
125
  ```
126
 
127
+ ## Evaluation
 
 
 
 
 
128
 
129
  ### Fresh 36-task mixed public holdout
130
 
131
+ A held-out slice of 18 FEVER verify-claim tasks plus 18 HotpotQA grounded-QA
132
+ tasks, drawn from public sources and de-duplicated against every training,
133
+ dev, and `probe_v0` row.
 
134
 
135
  | Stack | Task | Strict | Evidence F1 | Quote F1 |
136
  |---|---:|---:|---:|---:|
 
139
  | Bridge + `deterministic_v3` | 0.8611 | 0.5833 | 0.8815 | 0.8815 |
140
  | **Pilot 3 + `deterministic_v3`** | **0.8889** | **0.5833** | **0.9093** | **0.9093** |
141
 
142
+ The standalone release beats the earlier bridge model on task accuracy,
143
+ evidence F1, and quote F1 in both raw and normalized form, ties normalized
144
+ strict, and roughly doubles raw quote F1 at the model level.
 
145
 
146
  ### Fixed dev triage slice (21 tasks)
147
 
 
151
 
152
  ### Untouched 104-task Hotpot shadow slice
153
 
154
+ Pilot 3 raw improved quote-faithful behavior over the raw bridge model on this
155
+ slice, and pilot 3 + `deterministic_v3` matched bridge +
156
+ `deterministic_v3` at the system level. That surface remains a narrative
157
+ parity result because the report does not publish per-metric cells for it.
158
 
159
+ ## Release architecture
160
 
161
+ This project ends in two finished artifacts, not one:
162
 
163
+ 1. **Standalone model release** — this page. Pilot 3 is the strongest
164
+ version of the project's evidence-faithful behavior that moved into the
165
+ model itself, evaluated across multiple non-`probe_v0` surfaces.
166
+ 2. **Benchmark-facing hybrid stack** — bridge `checkpoint-2` plus the
167
+ `deterministic_v3` packet-local normalizer. That stack is the benchmark
168
+ winner and the only configuration that clears every gate on frozen held-out
169
+ `probe_v0`.
 
 
 
 
170
 
171
+ The separation is deliberate. This page is for the standalone release you can
172
+ download now. The benchmark winner is documented here because it explains the
173
+ project's full result, not because those perfect `probe_v0` numbers belong to
174
+ the adapter alone.
175
 
176
  ## Intended use
177
 
178
+ Use this release for work that has to stay inside a fixed body of text:
 
179
 
180
  - bounded document QA with explicit evidence requirements,
181
+ - claim verification and grounded QA from closed evidence packets,
182
+ - policy, compliance, contract, and internal-document workflows where each
183
+ answer must be justified from the provided text,
184
  - research on evidence-faithful reasoning and abstention behavior.
185
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
186
  ## Limitations
187
 
188
+ - The downloadable artifact is the LoRA adapter only. The base model is
189
+ required.
190
+ - The `deterministic_v3` packet-local normalizer is not included in this
191
+ download. The benchmark-winning configuration is adapter + normalizer, while
192
+ the adapter alone reproduces the standalone-model results shown above.
193
+ - Perfect `probe_v0` belongs to the benchmark-facing hybrid stack, not to this
194
+ adapter alone.
195
+ - Specialized for closed-packet reasoning, not open-ended chat or open-domain
196
+ QA.
197
+ - Frozen `probe_v0` item-level contents are intentionally not published with
198
+ the release.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
199
 
200
  ## Citation
201
 
202
+ References:
203
+
204
+ - Base model:
205
+ [Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2](https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2)
206
+ - Datasets:
207
+ [fever/fever](https://huggingface.co/datasets/fever/fever),
208
+ [hotpotqa/hotpot_qa](https://huggingface.co/datasets/hotpotqa/hotpot_qa)
209
+ - Technical brief (PDF):
210
+ [evidence_faithful_reasoning_release_brief.pdf](./evidence_faithful_reasoning_release_brief.pdf)
211
+ - Technical note:
212
+ [technical_note_evidence_faithful_reasoning.md](./technical_note_evidence_faithful_reasoning.md)
213
+
214
  ```bibtex
215
+ @misc{darcar0_evidence_faithful_reasoning_pilot_3_2026,
216
+ title = {Evidence-Faithful Reasoning: Pilot 3},
217
+ author = {darcar0},
218
  year = {2026},
219
  howpublished = {Hugging Face model release},
220
+ url = {https://huggingface.co/darcar0/evidence-faithful-reasoning-pilot-3}
 
221
  }
222
  ```