AnkitAI commited on
Commit
b25cef4
·
verified ·
1 Parent(s): 606c0ed

Add measured BFCL V3 AST results (base vs Parable, identical harness)

Browse files
Files changed (1) hide show
  1. README.md +25 -0
README.md CHANGED
@@ -122,6 +122,31 @@ With Claude Fable 5 now retired, genuine self-authored Fable traces are a fixed,
122
  - Not trained for: multi-file repo navigation, vision, non-English.
123
  - Inherits Granite-4.1-3B's knowledge cutoff. Treat generated commands as drafts to review.
124
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
125
  ## Base & license
126
 
127
  Weights: **Apache-2.0** (inherited from [ibm-granite/granite-4.1-3b](https://huggingface.co/ibm-granite/granite-4.1-3b)). Training data: Fable-5-traces **AGPL-3.0**, gpt5.5-terminal **MIT** — since traces originate from third-party assistants, their terms may apply to downstream training; check before commercial distillation.
 
122
  - Not trained for: multi-file repo navigation, vision, non-English.
123
  - Inherits Granite-4.1-3B's knowledge cutoff. Treat generated commands as drafts to review.
124
 
125
+ ## Evaluation
126
+
127
+ ### Function calling (BFCL V3, AST subset)
128
+
129
+ Measured 2026-07-29: bfcl-eval at gorilla main, prompting mode, Q4_K_M
130
+ GGUFs served by llama.cpp on a T4, base and Parable under the identical
131
+ harness. Categories: simple_python / multiple / parallel /
132
+ parallel_multiple (400/200/200/200 items). Raw generations and score
133
+ files: [parable-v2-artifacts](https://huggingface.co/AnkitAI/parable-v2-artifacts)
134
+ under `verify/bfcl/`.
135
+
136
+ | | simple_python | multiple | parallel | parallel_multiple |
137
+ |---|---|---|---|---|
138
+ | Granite-4.1-3B base | 0.848 | 0.790 | 0.710 | 0.665 |
139
+ | **This model (chat variant)** | 0.413 | 0.605 | 0.320 | 0.425 |
140
+
141
+ **Plain reading: do not use this variant for function calling.** The
142
+ drop is not random wrongness: sampled generations show the model
143
+ intermittently answering with args-only tool-call JSON (for example
144
+ `{"base": 10, "height": 5}`) instead of a function call, which the
145
+ AST scorer rejects. That is trace-scaffolding format bleeding into
146
+ standalone tasks, the failure mode the series paper names *session
147
+ leakage* (Section 6 of the report). Use the base model for tool-calling
148
+ workloads; use this variant for its reasoning prose.
149
+
150
  ## Base & license
151
 
152
  Weights: **Apache-2.0** (inherited from [ibm-granite/granite-4.1-3b](https://huggingface.co/ibm-granite/granite-4.1-3b)). Training data: Fable-5-traces **AGPL-3.0**, gpt5.5-terminal **MIT** — since traces originate from third-party assistants, their terms may apply to downstream training; check before commercial distillation.