darcar0 commited on
Commit
b314f67
·
verified ·
1 Parent(s): dab6650

Simplify public release surfaces around the technical note

Browse files
README.md CHANGED
@@ -26,12 +26,11 @@ tags:
26
  *The standalone model release from Evidence-Faithful Reasoning, built on the
27
  Qwen 3.5 Opus Distilled 27B base.*
28
 
29
- Quotebound 27B is the public release name for the standalone model
30
- previously referred to internally as `pilot 3`. I built it as the
31
- downloadable model release for Evidence-Faithful Reasoning: a LoRA adapter
32
- that turns its reasoning-distilled 27B base model into an evidence-first
33
- reader for closed packets of source text. Every answer has to land on the
34
- right evidence units, quote them verbatim, and stop with
35
  `Insufficient evidence.` when the packet does not justify a claim.
36
 
37
  ![Fresh public holdout: Quotebound 27B versus the prior bridge model](./standalone_holdout_comparison.svg)
@@ -59,10 +58,8 @@ quote normalizer carries the full stack to `0.9093` quote F1.*
59
 
60
  ## Read next
61
 
62
- - [Technical brief (PDF)](./evidence_faithful_reasoning_release_brief.pdf) — short, citation-friendly summary.
63
  - [Technical note](./technical_note_evidence_faithful_reasoning.md) — full method, results, and discussion.
64
  - [Frozen benchmark progression chart](./benchmark_progression.svg)
65
- - [Release architecture chart](./project_release_arc.svg)
66
 
67
  ## Quick start
68
 
@@ -225,8 +222,6 @@ Use this release for work that has to stay inside a fixed body of text:
225
  - Datasets:
226
  [fever/fever](https://huggingface.co/datasets/fever/fever),
227
  [hotpotqa/hotpot_qa](https://huggingface.co/datasets/hotpotqa/hotpot_qa)
228
- - Technical brief (PDF):
229
- [evidence_faithful_reasoning_release_brief.pdf](./evidence_faithful_reasoning_release_brief.pdf)
230
  - Technical note:
231
  [technical_note_evidence_faithful_reasoning.md](./technical_note_evidence_faithful_reasoning.md)
232
 
 
26
  *The standalone model release from Evidence-Faithful Reasoning, built on the
27
  Qwen 3.5 Opus Distilled 27B base.*
28
 
29
+ Quotebound 27B is the downloadable model release for
30
+ Evidence-Faithful Reasoning: a LoRA adapter that turns its
31
+ reasoning-distilled 27B base model into an evidence-first reader for
32
+ closed packets of source text. Every answer has to land on the right
33
+ evidence units, quote them verbatim, and stop with
 
34
  `Insufficient evidence.` when the packet does not justify a claim.
35
 
36
  ![Fresh public holdout: Quotebound 27B versus the prior bridge model](./standalone_holdout_comparison.svg)
 
58
 
59
  ## Read next
60
 
 
61
  - [Technical note](./technical_note_evidence_faithful_reasoning.md) — full method, results, and discussion.
62
  - [Frozen benchmark progression chart](./benchmark_progression.svg)
 
63
 
64
  ## Quick start
65
 
 
222
  - Datasets:
223
  [fever/fever](https://huggingface.co/datasets/fever/fever),
224
  [hotpotqa/hotpot_qa](https://huggingface.co/datasets/hotpotqa/hotpot_qa)
 
 
225
  - Technical note:
226
  [technical_note_evidence_faithful_reasoning.md](./technical_note_evidence_faithful_reasoning.md)
227
 
evidence_faithful_reasoning_release_brief.md CHANGED
@@ -13,7 +13,6 @@ Companion files:
13
  - [`technical_note_evidence_faithful_reasoning.md`](./technical_note_evidence_faithful_reasoning.md)
14
  - [`standalone_holdout_comparison.svg`](./standalone_holdout_comparison.svg)
15
  - [`benchmark_progression.svg`](./benchmark_progression.svg)
16
- - [`project_release_arc.svg`](./project_release_arc.svg)
17
 
18
  ## Executive summary
19
 
@@ -43,11 +42,10 @@ behavior at the model level.
43
 
44
  ## Quotebound 27B
45
 
46
- Quotebound 27B is the public release name for the standalone model
47
- previously referred to internally as `pilot 3`. It is the strongest
48
- standalone model the project produced and the artifact most readers will
49
- load first. It is the first standalone checkpoint in the project to hold
50
- up across multiple evaluation surfaces beyond the held-out probe.
51
 
52
  Fresh 36-task mixed public holdout:
53
 
@@ -80,19 +78,17 @@ remaining quote-faithful gap without leaving the closed-packet boundary.
80
  | Contrastive consistency | **1.0000** |
81
  | Invalid / missing rate | **0.0000** |
82
 
83
- ## Project arc and stopping point
84
 
85
  The release has two public faces: Quotebound 27B, the standalone model that
86
  loads directly from Hugging Face, and a benchmark-facing hybrid stack that
87
  closes the last quote-faithfulness gap on the frozen held-out probe. The
88
  split is part of the project story, not hidden behind the fine print.
89
 
90
- A narrow follow-up — internally `pilot 4` — fixed one specific FEVER
91
- month/date temporal-insufficiency case but weakened broader behavior on
92
- larger evaluation surfaces. The project read that outcome as a stop signal
93
- rather than running additional local fixes, and froze at the point where
94
- Quotebound 27B was strongest and the benchmark-facing system was already
95
- complete.
96
 
97
  ## Intended use and boundaries
98
 
@@ -121,5 +117,3 @@ Important boundaries:
121
  [`standalone_holdout_comparison.svg`](./standalone_holdout_comparison.svg)
122
  - Frozen benchmark progression chart:
123
  [`benchmark_progression.svg`](./benchmark_progression.svg)
124
- - Release architecture chart:
125
- [`project_release_arc.svg`](./project_release_arc.svg)
 
13
  - [`technical_note_evidence_faithful_reasoning.md`](./technical_note_evidence_faithful_reasoning.md)
14
  - [`standalone_holdout_comparison.svg`](./standalone_holdout_comparison.svg)
15
  - [`benchmark_progression.svg`](./benchmark_progression.svg)
 
16
 
17
  ## Executive summary
18
 
 
42
 
43
  ## Quotebound 27B
44
 
45
+ Quotebound 27B is the strongest standalone model the project produced and
46
+ the artifact most readers will load first. It is the first standalone
47
+ checkpoint in the project to hold up across multiple evaluation surfaces
48
+ beyond the held-out probe.
 
49
 
50
  Fresh 36-task mixed public holdout:
51
 
 
78
  | Contrastive consistency | **1.0000** |
79
  | Invalid / missing rate | **0.0000** |
80
 
81
+ ## Release boundary
82
 
83
  The release has two public faces: Quotebound 27B, the standalone model that
84
  loads directly from Hugging Face, and a benchmark-facing hybrid stack that
85
  closes the last quote-faithfulness gap on the frozen held-out probe. The
86
  split is part of the project story, not hidden behind the fine print.
87
 
88
+ The public release stops at the point where the strongest benchmark-facing
89
+ system and the strongest standalone model were both clearly in hand. That
90
+ keeps the package centered on finished artifacts rather than on local
91
+ variant history.
 
 
92
 
93
  ## Intended use and boundaries
94
 
 
117
  [`standalone_holdout_comparison.svg`](./standalone_holdout_comparison.svg)
118
  - Frozen benchmark progression chart:
119
  [`benchmark_progression.svg`](./benchmark_progression.svg)
 
 
evidence_faithful_reasoning_release_brief.pdf DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:4016a00664573e528b22b0d93a4b0d5ebe740c0c6f4a4e9fd6f6874a02c55d25
3
- size 243746
 
 
 
 
project_release_arc.svg DELETED
technical_note_evidence_faithful_reasoning.md CHANGED
@@ -13,14 +13,11 @@ packet units as evidence, quote those units verbatim, and abstain with
13
  project ends in two finished results from the same frame. The first is a
14
  hybrid system — a trained bridge checkpoint plus a packet-local quote
15
  normalizer — that clears every gate of the contract on the frozen
16
- held-out probe (`probe_v0`). The second is Quotebound 27B, the public
17
- release name for the standalone model previously referred to internally as
18
- `pilot 3`, released as a LoRA adapter on Hugging Face
19
  ([`darcar0/quotebound-27b`](https://huggingface.co/darcar0/quotebound-27b)).
20
- A narrow follow-up — internally `pilot 4` — was rejected as a stop signal
21
- after it traded one local fix for broader regressions. The hybrid system
22
- is the benchmark-facing release; Quotebound 27B is the downloadable model
23
- release.
24
 
25
  **Keywords:** evidence-faithful reasoning, grounded QA, claim
26
  verification, closed-packet reasoning, abstention, attribution.
@@ -77,8 +74,8 @@ normalizer was layered on as a finishing repair, and held-out evaluation
77
  was rerun on the protected split. Once the hybrid stack solved the
78
  held-out probe, a teacher–student distillation cycle pushed the winning
79
  behavior back into the model itself, evaluated entirely off the held-out
80
- probe. The cycle was stopped when the strongest standalone release
81
- artifact emerged and a focused follow-up showed clear regressions.
82
 
83
  Train and dev surfaces are derived from public FEVER-style
84
  verify-claim data, public HotpotQA-style grounded-QA data, and
@@ -126,14 +123,10 @@ question inside the same frame: how much of that winning behavior can be
126
  moved into the model itself, evaluated on surfaces outside the held-out
127
  probe?
128
 
129
- That question produced Quotebound 27B — the public release name for the
130
- standalone model previously referred to internally as `pilot 3` — a
131
- teacher-student distillation checkpoint published as a LoRA adapter on top
132
- of
133
  [`Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2`](https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2).
134
 
135
- - **Internal checkpoint:**
136
- `outputs/sft_v1_v2_teacher_distill_pilot_v3_partialdev/checkpoint-16`
137
  - **Public release identity:**
138
  [`darcar0/quotebound-27b`](https://huggingface.co/darcar0/quotebound-27b)
139
 
@@ -182,18 +175,13 @@ at the system level on this slice. The standalone freeze memo
182
  outcome; per-metric numbers were not recorded for this surface, so it stands
183
  as a narrative parity result rather than a table cell.
184
 
185
- ## 8. Stop-signal follow-up
186
-
187
- A narrow follow-up — internally `pilot 4` — was built to fix one specific
188
- FEVER month/date temporal-insufficiency error. It fixed that single row but
189
- weakened broader behavior on the larger evaluation surfaces. The project read
190
- that outcome as a stop signal: further local-fix iteration was trading visible
191
- gains for wider regressions, so Quotebound 27B froze at the prior
192
- checkpoint. Canonical memo:
193
- `reports/sft_v1_v2_teacher_distill_pilot_v4_partialdev_status.md`.
194
 
195
- See [Figure 3: project release arc](./project_release_arc.svg) for the full
196
- arc from baseline through stop signal.
 
 
 
197
 
198
  ## 9. Discussion
199
 
@@ -216,11 +204,10 @@ model in the project to hold up across multiple evaluation surfaces beyond
216
  the held-out probe; calling it the standalone release artifact is faithful
217
  to that evidence.
218
 
219
- **Why the stop signal matters.** The follow-up was a deliberate narrow
220
- refinement. It worked on the one row it targeted and regressed broader
221
- behavior. Reading that result as a stop signal — rather than running
222
- additional local fixes — is the discipline the project decided to keep
223
- visible in the release trail.
224
 
225
  **Distinction held throughout.** Perfect frozen `probe_v0` belongs to the
226
  hybrid stack, not to the standalone adapter alone. The release ships both
@@ -241,8 +228,9 @@ results without collapsing them into one claim.
241
  the project to hold up across multiple evaluation surfaces beyond
242
  `probe_v0`, and that roughly doubles raw quote-faithful behavior over
243
  the earlier bridge.
244
- 5. A documented stop signal that marks the point at which further local-fix
245
- iteration began trading visible gains for wider regressions.
 
246
 
247
  ## 11. Intended use and limitations
248
 
@@ -274,7 +262,6 @@ Canonical project surfaces:
274
  - final artifact memo: `reports/sft_v1_final_artifact_status.md`
275
  - standalone freeze memo: `reports/standalone_model_v2_freeze_memo.md`
276
  - fresh holdout comparison: `reports/standalone_model_v2_holdout_v1_bridge_vs_pilot3_status.md`
277
- - stop-signal memo: `reports/sft_v1_v2_teacher_distill_pilot_v4_partialdev_status.md`
278
- - release brief: `evidence_faithful_reasoning_release_brief.pdf`
279
  - model card: `README.md`
 
280
  - Hugging Face release: [`darcar0/quotebound-27b`](https://huggingface.co/darcar0/quotebound-27b)
 
13
  project ends in two finished results from the same frame. The first is a
14
  hybrid system — a trained bridge checkpoint plus a packet-local quote
15
  normalizer — that clears every gate of the contract on the frozen
16
+ held-out probe (`probe_v0`). The second is Quotebound 27B, released as a
17
+ LoRA adapter on Hugging Face
 
18
  ([`darcar0/quotebound-27b`](https://huggingface.co/darcar0/quotebound-27b)).
19
+ The hybrid system is the benchmark-facing release; Quotebound 27B is the
20
+ downloadable model release.
 
 
21
 
22
  **Keywords:** evidence-faithful reasoning, grounded QA, claim
23
  verification, closed-packet reasoning, abstention, attribution.
 
74
  was rerun on the protected split. Once the hybrid stack solved the
75
  held-out probe, a teacher–student distillation cycle pushed the winning
76
  behavior back into the model itself, evaluated entirely off the held-out
77
+ probe. The public release stops at the strongest standalone artifact that
78
+ preserved the broader gains reported below.
79
 
80
  Train and dev surfaces are derived from public FEVER-style
81
  verify-claim data, public HotpotQA-style grounded-QA data, and
 
123
  moved into the model itself, evaluated on surfaces outside the held-out
124
  probe?
125
 
126
+ That question produced Quotebound 27B, a teacher-student distillation
127
+ checkpoint published as a LoRA adapter on top of
 
 
128
  [`Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2`](https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2).
129
 
 
 
130
  - **Public release identity:**
131
  [`darcar0/quotebound-27b`](https://huggingface.co/darcar0/quotebound-27b)
132
 
 
175
  outcome; per-metric numbers were not recorded for this surface, so it stands
176
  as a narrative parity result rather than a table cell.
177
 
178
+ ## 8. Release boundary
 
 
 
 
 
 
 
 
179
 
180
+ The public release centers on the two strongest finished artifacts from the
181
+ project: the benchmark-winning hybrid stack on frozen `probe_v0`, and
182
+ Quotebound 27B as the strongest standalone model. That boundary keeps the
183
+ package focused on the results that held up across the evaluation surfaces
184
+ reported here.
185
 
186
  ## 9. Discussion
187
 
 
204
  the held-out probe; calling it the standalone release artifact is faithful
205
  to that evidence.
206
 
207
+ **Why the release stops here.** The package is intentionally centered on
208
+ the strongest benchmark-facing system and the strongest standalone model.
209
+ That keeps the public release focused on finished results rather than on
210
+ intermediate local variations.
 
211
 
212
  **Distinction held throughout.** Perfect frozen `probe_v0` belongs to the
213
  hybrid stack, not to the standalone adapter alone. The release ships both
 
228
  the project to hold up across multiple evaluation surfaces beyond
229
  `probe_v0`, and that roughly doubles raw quote-faithful behavior over
230
  the earlier bridge.
231
+ 5. A public artifact set centered on the benchmark-winning hybrid system,
232
+ Quotebound 27B, and the technical note that explains how the two fit
233
+ together.
234
 
235
  ## 11. Intended use and limitations
236
 
 
262
  - final artifact memo: `reports/sft_v1_final_artifact_status.md`
263
  - standalone freeze memo: `reports/standalone_model_v2_freeze_memo.md`
264
  - fresh holdout comparison: `reports/standalone_model_v2_holdout_v1_bridge_vs_pilot3_status.md`
 
 
265
  - model card: `README.md`
266
+ - technical note: `technical_note_evidence_faithful_reasoning.md`
267
  - Hugging Face release: [`darcar0/quotebound-27b`](https://huggingface.co/darcar0/quotebound-27b)