darcar0 commited on
Commit
5fb78cb
·
verified ·
1 Parent(s): e558a1c

Initial private pilot 3 adapter upload

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,220 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ license: mit
5
+ base_model:
6
+ - Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2
7
+ library_name: transformers
8
+ pipeline_tag: text-generation
9
+ tags:
10
+ - reasoning
11
+ - evidence-grounding
12
+ - attribution
13
+ - fever
14
+ - hotpotqa
15
+ - lora
16
+ - peft
17
+ - distillation
18
+ - research
19
+ ---
20
+
21
+ # Evidence-Faithful Reasoning Pilot 3
22
+
23
+ This release is the standalone-model-v2 checkpoint from the
24
+ evidence-faithful reasoning project.
25
+
26
+ The project studies a stricter target than ordinary question answering. The
27
+ system must answer from a bounded packet of source text, identify the right
28
+ evidence, quote exact supporting text, and abstain with
29
+ `Insufficient evidence.` when the packet does not justify a claim.
30
+
31
+ Pilot 3 is the strongest standalone model artifact produced by the project. It
32
+ is the model this release should lead with for direct download and use.
33
+
34
+ ## What This Release Is
35
+
36
+ This card is for the pilot 3 standalone adapter release:
37
+
38
+ - checkpoint:
39
+ `outputs/sft_v1_v2_teacher_distill_pilot_v3_partialdev/checkpoint-16`
40
+
41
+ Pilot 3 packages the strongest version of the project’s evidence-faithful
42
+ behavior that successfully moved into one model. The same project also
43
+ produced a benchmark-winning hybrid system for the frozen held-out benchmark,
44
+ but this page is for the main downloadable model artifact.
45
+
46
+ ## Why It Matters
47
+
48
+ Pilot 3 is important for three reasons:
49
+
50
+ 1. it is the first standalone model in the project that stays strong across
51
+ multiple non-`probe_v0` evaluation surfaces
52
+ 2. it materially improves raw quote-faithful behavior over the earlier bridge
53
+ model
54
+ 3. it gives the project a clean downloadable model artifact, not only a
55
+ benchmark-facing system result
56
+
57
+ ## At A Glance
58
+
59
+ - model role:
60
+ main downloadable model from the project
61
+ - benchmark-facing winner:
62
+ bridge checkpoint-2 + `deterministic_v3`
63
+ - public model story:
64
+ pilot 3 is the strongest direct model artifact; the hybrid stack remains the
65
+ strongest full benchmark system
66
+
67
+ ## Quick Start
68
+
69
+ Pilot 3 is a LoRA adapter on top of the base model listed above. To use it as
70
+ the direct standalone release, load the base model and attach the published
71
+ pilot 3 adapter:
72
+
73
+ ```python
74
+ from peft import PeftModel
75
+ from transformers import AutoModelForCausalLM, AutoTokenizer
76
+
77
+ base_id = "Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2"
78
+ adapter_id = "darcar0/evidence-faithful-reasoning-pilot-3"
79
+
80
+ tokenizer = AutoTokenizer.from_pretrained(base_id)
81
+ base = AutoModelForCausalLM.from_pretrained(base_id, device_map="auto")
82
+ model = PeftModel.from_pretrained(base, adapter_id)
83
+ ```
84
+
85
+ If you want to reproduce the benchmark-winning hybrid result instead, use the
86
+ project repository’s bridge checkpoint plus `deterministic_v3`.
87
+
88
+ ## Strongest Public Evidence
89
+
90
+ Pilot 3 holds up across all three non-`probe_v0` surfaces used for standalone
91
+ selection:
92
+
93
+ ### Fixed dev triage slice
94
+
95
+ Pilot 3 + `deterministic_v3`:
96
+
97
+ - task `1.0000`
98
+ - strict `0.6190`
99
+ - evidence F1 `0.8320`
100
+ - quote F1 `0.7095`
101
+
102
+ ### Untouched 104-task Hotpot shadow slice
103
+
104
+ - pilot 3 raw materially improved quote-faithful behavior over raw bridge
105
+ - pilot 3 + `deterministic_v3` matched bridge + `deterministic_v3`
106
+
107
+ ### Fresh 36-task mixed public holdout
108
+
109
+ Bridge + `deterministic_v3`:
110
+
111
+ - task `0.8611`
112
+ - strict `0.5833`
113
+ - evidence F1 `0.8815`
114
+ - quote F1 `0.8815`
115
+
116
+ Pilot 3 + `deterministic_v3`:
117
+
118
+ - task `0.8889`
119
+ - strict `0.5833`
120
+ - evidence F1 `0.9093`
121
+ - quote F1 `0.9093`
122
+
123
+ Raw bridge vs raw pilot 3 on the same holdout:
124
+
125
+ - bridge raw
126
+ - task `0.8611`
127
+ - strict `0.2222`
128
+ - evidence F1 `0.8815`
129
+ - quote F1 `0.3343`
130
+ - pilot 3 raw
131
+ - task `0.8889`
132
+ - strict `0.4444`
133
+ - evidence F1 `0.9093`
134
+ - quote F1 `0.6815`
135
+
136
+ This means pilot 3 is both the cleanest release artifact from the project and
137
+ a materially stronger standalone model than the earlier bridge checkpoint on a
138
+ fresh public holdout.
139
+
140
+ ## How This Fits Into The Full Project
141
+
142
+ The project has two finished outcomes:
143
+
144
+ 1. **Benchmark-facing winner**
145
+ - bridge checkpoint-2 + `deterministic_v3`
146
+ - perfect on frozen held-out `probe_v0`
147
+ 2. **Main downloadable model**
148
+ - pilot 3
149
+ - strongest standalone model released from the project
150
+
151
+ That distinction is deliberate.
152
+
153
+ The hybrid stack is the strongest full benchmark system. Pilot 3 is the
154
+ strongest clean model artifact.
155
+
156
+ ## Intended Use
157
+
158
+ This release is intended for:
159
+
160
+ - research on evidence-faithful reasoning
161
+ - bounded document QA with explicit evidence requirements
162
+ - claim verification and grounded QA from fixed evidence packets
163
+ - policy, compliance, contract, or internal-document reasoning workflows where
164
+ answers must be justified from a fixed packet of text
165
+
166
+ This is a specialized grounded reasoning model, not a general-purpose chatbot
167
+ replacement.
168
+
169
+ ## Training / Data Notes
170
+
171
+ The training and evaluation surfaces are public-data-backed and derived from:
172
+
173
+ - FEVER-style verify-claim data
174
+ - HotpotQA-style grounded QA data
175
+ - project-local bounded packet scaffolds
176
+
177
+ The held-out benchmark `probe_v0` remained frozen and was **not** used as a
178
+ tuning surface for this standalone selection cycle.
179
+
180
+ ## Why Pilot 3 Was Frozen
181
+
182
+ Pilot 4 was a deliberately narrow refinement meant to fix one FEVER
183
+ month/date-style temporal insufficiency error.
184
+
185
+ It fixed that targeted row.
186
+
187
+ But it also weakened broader behavior on the larger evaluation surfaces.
188
+
189
+ That made pilot 4 a stop signal, not a better release.
190
+
191
+ So pilot 3 is the checkpoint the project froze as the strongest standalone
192
+ model release.
193
+
194
+ ## Limitations
195
+
196
+ - Pilot 3 is the main downloadable model from the project. The separate
197
+ benchmark-facing winner is bridge checkpoint-2 + `deterministic_v3`.
198
+ - Perfect `probe_v0` belongs to the hybrid stack, not to pilot 3 alone.
199
+ - Perfect `probe_v0` should not be interpreted as proof of general faithful
200
+ reasoning.
201
+ - This release is specialized around bounded evidence packets, not general
202
+ chat.
203
+
204
+ ## Repository Surfaces
205
+
206
+ Canonical project surfaces:
207
+
208
+ - standalone freeze memo:
209
+ `reports/standalone_model_v2_freeze_memo.md`
210
+ - fresh standalone holdout comparison:
211
+ `reports/standalone_model_v2_holdout_v1_bridge_vs_pilot3_status.md`
212
+ - benchmark-facing final artifact:
213
+ `reports/sft_v1_final_artifact_status.md`
214
+ - release page:
215
+ `release/index.html`
216
+
217
+ ## Citation
218
+
219
+ If you use this release, cite the project repository and the standalone freeze
220
+ memo.
adapter_config.json ADDED
@@ -0,0 +1,46 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": null,
18
+ "loftq_config": {},
19
+ "lora_alpha": 32,
20
+ "lora_bias": false,
21
+ "lora_dropout": 0.1,
22
+ "megatron_config": null,
23
+ "megatron_core": "megatron.core",
24
+ "modules_to_save": null,
25
+ "peft_type": "LORA",
26
+ "peft_version": "0.18.1",
27
+ "qalora_group_size": 16,
28
+ "r": 16,
29
+ "rank_pattern": {},
30
+ "revision": null,
31
+ "target_modules": [
32
+ "o_proj",
33
+ "k_proj",
34
+ "down_proj",
35
+ "gate_proj",
36
+ "v_proj",
37
+ "up_proj",
38
+ "q_proj"
39
+ ],
40
+ "target_parameters": null,
41
+ "task_type": "CAUSAL_LM",
42
+ "trainable_token_indices": null,
43
+ "use_dora": false,
44
+ "use_qalora": false,
45
+ "use_rslora": false
46
+ }
adapter_model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:57479e3f6df4b0838dcb5615ba8e10253cdbd4cecdfe193b8bf0fca3caa85d04
3
+ size 159452256
chat_template.jinja ADDED
@@ -0,0 +1,154 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- set image_count = namespace(value=0) %}
2
+ {%- set video_count = namespace(value=0) %}
3
+ {%- macro render_content(content, do_vision_count, is_system_content=false) %}
4
+ {%- if content is string %}
5
+ {{- content }}
6
+ {%- elif content is iterable and content is not mapping %}
7
+ {%- for item in content %}
8
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
9
+ {%- if is_system_content %}
10
+ {{- raise_exception('System message cannot contain images.') }}
11
+ {%- endif %}
12
+ {%- if do_vision_count %}
13
+ {%- set image_count.value = image_count.value + 1 %}
14
+ {%- endif %}
15
+ {%- if add_vision_id %}
16
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
17
+ {%- endif %}
18
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
19
+ {%- elif 'video' in item or item.type == 'video' %}
20
+ {%- if is_system_content %}
21
+ {{- raise_exception('System message cannot contain videos.') }}
22
+ {%- endif %}
23
+ {%- if do_vision_count %}
24
+ {%- set video_count.value = video_count.value + 1 %}
25
+ {%- endif %}
26
+ {%- if add_vision_id %}
27
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
28
+ {%- endif %}
29
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
30
+ {%- elif 'text' in item %}
31
+ {{- item.text }}
32
+ {%- else %}
33
+ {{- raise_exception('Unexpected item type in content.') }}
34
+ {%- endif %}
35
+ {%- endfor %}
36
+ {%- elif content is none or content is undefined %}
37
+ {{- '' }}
38
+ {%- else %}
39
+ {{- raise_exception('Unexpected content type.') }}
40
+ {%- endif %}
41
+ {%- endmacro %}
42
+ {%- if not messages %}
43
+ {{- raise_exception('No messages provided.') }}
44
+ {%- endif %}
45
+ {%- if tools and tools is iterable and tools is not mapping %}
46
+ {{- '<|im_start|>system\n' }}
47
+ {{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
48
+ {%- for tool in tools %}
49
+ {{- "\n" }}
50
+ {{- tool | tojson }}
51
+ {%- endfor %}
52
+ {{- "\n</tools>" }}
53
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
54
+ {%- if messages[0].role == 'system' %}
55
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
56
+ {%- if content %}
57
+ {{- '\n\n' + content }}
58
+ {%- endif %}
59
+ {%- endif %}
60
+ {{- '<|im_end|>\n' }}
61
+ {%- else %}
62
+ {%- if messages[0].role == 'system' %}
63
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
64
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
65
+ {%- endif %}
66
+ {%- endif %}
67
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
68
+ {%- for message in messages[::-1] %}
69
+ {%- set index = (messages|length - 1) - loop.index0 %}
70
+ {%- if ns.multi_step_tool and message.role == "user" %}
71
+ {%- set content = render_content(message.content, false)|trim %}
72
+ {%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
73
+ {%- set ns.multi_step_tool = false %}
74
+ {%- set ns.last_query_index = index %}
75
+ {%- endif %}
76
+ {%- endif %}
77
+ {%- endfor %}
78
+ {%- if ns.multi_step_tool %}
79
+ {{- raise_exception('No user query found in messages.') }}
80
+ {%- endif %}
81
+ {%- for message in messages %}
82
+ {%- set content = render_content(message.content, true)|trim %}
83
+ {%- if message.role == "system" %}
84
+ {%- if not loop.first %}
85
+ {{- raise_exception('System message must be at the beginning.') }}
86
+ {%- endif %}
87
+ {%- elif message.role == "user" %}
88
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
89
+ {%- elif message.role == "assistant" %}
90
+ {%- set reasoning_content = '' %}
91
+ {%- if message.reasoning_content is string %}
92
+ {%- set reasoning_content = message.reasoning_content %}
93
+ {%- else %}
94
+ {%- if '</think>' in content %}
95
+ {%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
96
+ {%- set content = content.split('</think>')[-1].lstrip('\n') %}
97
+ {%- endif %}
98
+ {%- endif %}
99
+ {%- set reasoning_content = reasoning_content|trim %}
100
+ {%- if loop.index0 > ns.last_query_index %}
101
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
102
+ {%- else %}
103
+ {{- '<|im_start|>' + message.role + '\n' + content }}
104
+ {%- endif %}
105
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
106
+ {%- for tool_call in message.tool_calls %}
107
+ {%- if tool_call.function is defined %}
108
+ {%- set tool_call = tool_call.function %}
109
+ {%- endif %}
110
+ {%- if loop.first %}
111
+ {%- if content|trim %}
112
+ {{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
113
+ {%- else %}
114
+ {{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
115
+ {%- endif %}
116
+ {%- else %}
117
+ {{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
118
+ {%- endif %}
119
+ {%- if tool_call.arguments is defined %}
120
+ {%- for args_name, args_value in tool_call.arguments|items %}
121
+ {{- '<parameter=' + args_name + '>\n' }}
122
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
123
+ {{- args_value }}
124
+ {{- '\n</parameter>\n' }}
125
+ {%- endfor %}
126
+ {%- endif %}
127
+ {{- '</function>\n</tool_call>' }}
128
+ {%- endfor %}
129
+ {%- endif %}
130
+ {{- '<|im_end|>\n' }}
131
+ {%- elif message.role == "tool" %}
132
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
133
+ {{- '<|im_start|>user' }}
134
+ {%- endif %}
135
+ {{- '\n<tool_response>\n' }}
136
+ {{- content }}
137
+ {{- '\n</tool_response>' }}
138
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
139
+ {{- '<|im_end|>\n' }}
140
+ {%- elif loop.last %}
141
+ {{- '<|im_end|>\n' }}
142
+ {%- endif %}
143
+ {%- else %}
144
+ {{- raise_exception('Unexpected message role.') }}
145
+ {%- endif %}
146
+ {%- endfor %}
147
+ {%- if add_generation_prompt %}
148
+ {{- '<|im_start|>assistant\n' }}
149
+ {%- if enable_thinking is defined and enable_thinking is false %}
150
+ {{- '<think>\n\n</think>\n\n' }}
151
+ {%- else %}
152
+ {{- '<think>\n' }}
153
+ {%- endif %}
154
+ {%- endif %}
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:87a7830d63fcf43bf241c3c5242e96e62dd3fdc29224ca26fed8ea333db72de4
3
+ size 19989343
tokenizer_config.json ADDED
@@ -0,0 +1,31 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "audio_bos_token": "<|audio_start|>",
4
+ "audio_eos_token": "<|audio_end|>",
5
+ "audio_token": "<|audio_pad|>",
6
+ "backend": "tokenizers",
7
+ "bos_token": null,
8
+ "clean_up_tokenization_spaces": false,
9
+ "eos_token": "<|im_end|>",
10
+ "errors": "replace",
11
+ "image_token": "<|image_pad|>",
12
+ "is_local": false,
13
+ "model_max_length": 262144,
14
+ "model_specific_special_tokens": {
15
+ "audio_bos_token": "<|audio_start|>",
16
+ "audio_eos_token": "<|audio_end|>",
17
+ "audio_token": "<|audio_pad|>",
18
+ "image_token": "<|image_pad|>",
19
+ "video_token": "<|video_pad|>",
20
+ "vision_bos_token": "<|vision_start|>",
21
+ "vision_eos_token": "<|vision_end|>"
22
+ },
23
+ "pad_token": "<|endoftext|>",
24
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
25
+ "split_special_tokens": false,
26
+ "tokenizer_class": "TokenizersBackend",
27
+ "unk_token": null,
28
+ "video_token": "<|video_pad|>",
29
+ "vision_bos_token": "<|vision_start|>",
30
+ "vision_eos_token": "<|vision_end|>"
31
+ }