AnkitAI commited on
Commit
f753211
·
verified ·
1 Parent(s): b25cef4

Reframe evaluation guidance: same facts, buyer guidance instead of verdicts

Browse files
Files changed (1) hide show
  1. README.md +4 -4
README.md CHANGED
@@ -138,14 +138,14 @@ under `verify/bfcl/`.
138
  | Granite-4.1-3B base | 0.848 | 0.790 | 0.710 | 0.665 |
139
  | **This model (chat variant)** | 0.413 | 0.605 | 0.320 | 0.425 |
140
 
141
- **Plain reading: do not use this variant for function calling.** The
142
- drop is not random wrongness: sampled generations show the model
143
  intermittently answering with args-only tool-call JSON (for example
144
  `{"base": 10, "height": 5}`) instead of a function call, which the
145
  AST scorer rejects. That is trace-scaffolding format bleeding into
146
  standalone tasks, the failure mode the series paper names *session
147
- leakage* (Section 6 of the report). Use the base model for tool-calling
148
- workloads; use this variant for its reasoning prose.
149
 
150
  ## Base & license
151
 
 
138
  | Granite-4.1-3B base | 0.848 | 0.790 | 0.710 | 0.665 |
139
  | **This model (chat variant)** | 0.413 | 0.605 | 0.320 | 0.425 |
140
 
141
+ For tool-calling workloads, use the base model; this variant is built
142
+ for reasoning prose. The drop has a specific mechanism: sampled generations show the model
143
  intermittently answering with args-only tool-call JSON (for example
144
  `{"base": 10, "height": 5}`) instead of a function call, which the
145
  AST scorer rejects. That is trace-scaffolding format bleeding into
146
  standalone tasks, the failure mode the series paper names *session
147
+ leakage* (Section 6 of the report). The reasoning-voice strengths this variant trains for are unaffected
148
+ on prose tasks.
149
 
150
  ## Base & license
151