darcar0 commited on
Commit
0cdc637
·
verified ·
1 Parent(s): e5e90d1

Upload evidence_faithful_reasoning_release_brief.md with huggingface_hub

Browse files
evidence_faithful_reasoning_release_brief.md CHANGED
@@ -8,15 +8,20 @@ Author: darcar0
8
  Hugging Face model release:
9
  [`darcar0/evidence-faithful-reasoning-pilot-3`](https://huggingface.co/darcar0/evidence-faithful-reasoning-pilot-3)
10
 
11
- Public GitHub release repo: `PUBLIC_REPO_URL`
 
 
 
 
 
12
 
13
  ## Executive summary
14
 
15
- I built this project because I wanted a model release where reasoning had to
16
- cash out into recoverable evidence instead of hiding behind fluent language.
17
- The result is a strict evidence-faithful reasoning benchmark, a
18
- benchmark-winning hybrid system, and pilot 3: the strongest standalone model
19
- from the same project, released as a LoRA adapter on Hugging Face.
20
 
21
  The contract is strict. On every bounded evidence packet, the system has to:
22
 
@@ -26,7 +31,7 @@ The contract is strict. On every bounded evidence packet, the system has to:
26
  4. abstain with `Insufficient evidence.` when the packet does not justify a
27
  claim.
28
 
29
- This release ships two finished outputs. The benchmark-facing winner is a
30
  hybrid stack — bridge `checkpoint-2` plus `deterministic_v3` packet-local
31
  quote normalization — that clears every gate on the frozen held-out
32
  `probe_v0` benchmark. The main downloadable artifact is pilot 3, the
@@ -34,33 +39,11 @@ standalone model release, which beats the earlier bridge model on a fresh
34
  mixed public holdout and roughly doubles raw quote-faithful behavior at the
35
  model level.
36
 
37
- ## Benchmark-facing winner
38
-
39
- The benchmark-facing winner is the strongest full system from the project.
40
- Training moved the model past the older frozen baseline. Deterministic
41
- packet-local normalization closed the remaining quote-faithfulness and strict
42
- grounded-success gap without leaving the bounded-packet setting.
43
-
44
- | Metric | Frozen probe_v0 |
45
- |---|---:|
46
- | Task success | **1.0000** |
47
- | Strict grounded success | **1.0000** |
48
- | Mean evidence F1 | **1.0000** |
49
- | Mean quote F1 | **1.0000** |
50
- | Verify label accuracy | **1.0000** |
51
- | Grounded QA accuracy | **1.0000** |
52
- | Contrastive consistency | **1.0000** |
53
- | Invalid / missing rate | **0.0000** |
54
-
55
- Canonical memo:
56
- [`reports/sft_v1_final_artifact_status.md`](../reports/sft_v1_final_artifact_status.md)
57
-
58
- ## Standalone model release: pilot 3
59
 
60
  Pilot 3 is the strongest standalone model from the project and the release I
61
- want people to download first. It is the first standalone checkpoint in the
62
- project to hold up across multiple non-`probe_v0` evaluation surfaces, and it
63
- moves a meaningful amount of the winning behavior into the model itself.
64
 
65
  Fresh 36-task mixed public holdout:
66
 
@@ -75,16 +58,30 @@ Pilot 3 beats bridge on task accuracy, evidence F1, and quote F1 in both raw
75
  and normalized form, ties normalized strict, and roughly doubles raw quote F1
76
  (`0.3343` → `0.6815`) at the model level.
77
 
78
- Canonical memos:
79
- - [`reports/standalone_model_v2_freeze_memo.md`](../reports/standalone_model_v2_freeze_memo.md)
80
- - [`reports/standalone_model_v2_holdout_v1_bridge_vs_pilot3_status.md`](../reports/standalone_model_v2_holdout_v1_bridge_vs_pilot3_status.md)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
81
 
82
  ## Project arc and stopping point
83
 
84
- The project followed a full research-engineering loop: baseline mapping,
85
- benchmark design, prompt and structure interventions, a training-backed bridge
86
- model, deterministic quote normalization, and then a teacher-student
87
- distillation cycle to move the winning behavior into the model itself.
88
 
89
  A targeted follow-up, pilot 4, fixed one specific FEVER
90
  month/date temporal-insufficiency case but weakened broader behavior on larger
@@ -113,6 +110,9 @@ Important boundaries:
113
  [`darcar0/evidence-faithful-reasoning-pilot-3`](https://huggingface.co/darcar0/evidence-faithful-reasoning-pilot-3)
114
  - Technical note:
115
  [`docs/technical_note_evidence_faithful_reasoning.md`](./technical_note_evidence_faithful_reasoning.md)
116
- - Release page:
117
- [`release/index.html`](../release/index.html)
118
- - Public GitHub release repo: `PUBLIC_REPO_URL`
 
 
 
 
8
  Hugging Face model release:
9
  [`darcar0/evidence-faithful-reasoning-pilot-3`](https://huggingface.co/darcar0/evidence-faithful-reasoning-pilot-3)
10
 
11
+ Companion files on this page:
12
+
13
+ - [`technical_note_evidence_faithful_reasoning.md`](./technical_note_evidence_faithful_reasoning.md)
14
+ - [`standalone_holdout_comparison.svg`](./standalone_holdout_comparison.svg)
15
+ - [`benchmark_progression.svg`](./benchmark_progression.svg)
16
+ - [`project_release_arc.svg`](./project_release_arc.svg)
17
 
18
  ## Executive summary
19
 
20
+ I built this project because I wanted a release where reasoning had to prove
21
+ itself from the evidence instead of hiding behind fluent language. The result
22
+ is a strict evidence-faithful benchmark, a benchmark-winning hybrid system,
23
+ and pilot 3: the standalone model release, shipped here as a LoRA adapter on
24
+ Hugging Face.
25
 
26
  The contract is strict. On every bounded evidence packet, the system has to:
27
 
 
31
  4. abstain with `Insufficient evidence.` when the packet does not justify a
32
  claim.
33
 
34
+ This release ends in two finished outputs. The benchmark-facing winner is a
35
  hybrid stack — bridge `checkpoint-2` plus `deterministic_v3` packet-local
36
  quote normalization — that clears every gate on the frozen held-out
37
  `probe_v0` benchmark. The main downloadable artifact is pilot 3, the
 
39
  mixed public holdout and roughly doubles raw quote-faithful behavior at the
40
  model level.
41
 
42
+ ## Standalone model release
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
43
 
44
  Pilot 3 is the strongest standalone model from the project and the release I
45
+ want people to open first. It is the first standalone checkpoint in the
46
+ project that holds up across multiple non-`probe_v0` evaluation surfaces.
 
47
 
48
  Fresh 36-task mixed public holdout:
49
 
 
58
  and normalized form, ties normalized strict, and roughly doubles raw quote F1
59
  (`0.3343` → `0.6815`) at the model level.
60
 
61
+ ## Benchmark-facing winner
62
+
63
+ The benchmark-facing hybrid stack is the strongest full system from the
64
+ project. Training moved the model past the older frozen baseline.
65
+ Deterministic packet-local normalization closed the remaining quote-faithful
66
+ gap without leaving the bounded-packet setting.
67
+
68
+ | Metric | Frozen probe_v0 |
69
+ |---|---:|
70
+ | Task success | **1.0000** |
71
+ | Strict grounded success | **1.0000** |
72
+ | Mean evidence F1 | **1.0000** |
73
+ | Mean quote F1 | **1.0000** |
74
+ | Verify label accuracy | **1.0000** |
75
+ | Grounded QA accuracy | **1.0000** |
76
+ | Contrastive consistency | **1.0000** |
77
+ | Invalid / missing rate | **0.0000** |
78
 
79
  ## Project arc and stopping point
80
 
81
+ The release has two public faces: a standalone model you can load directly and
82
+ a benchmark-facing hybrid stack that closes the last quote-faithfulness gap.
83
+ That split is part of the project story, not something hidden behind the fine
84
+ print.
85
 
86
  A targeted follow-up, pilot 4, fixed one specific FEVER
87
  month/date temporal-insufficiency case but weakened broader behavior on larger
 
110
  [`darcar0/evidence-faithful-reasoning-pilot-3`](https://huggingface.co/darcar0/evidence-faithful-reasoning-pilot-3)
111
  - Technical note:
112
  [`docs/technical_note_evidence_faithful_reasoning.md`](./technical_note_evidence_faithful_reasoning.md)
113
+ - Fresh public holdout chart:
114
+ [`standalone_holdout_comparison.svg`](./standalone_holdout_comparison.svg)
115
+ - Frozen benchmark progression chart:
116
+ [`benchmark_progression.svg`](./benchmark_progression.svg)
117
+ - Release architecture chart:
118
+ [`project_release_arc.svg`](./project_release_arc.svg)