AnkitAI commited on
Commit
86acda1
·
verified ·
1 Parent(s): 92a8e85

Add measured BFCL V3 AST results (base vs Parable, identical harness)

Browse files
Files changed (1) hide show
  1. README.md +20 -0
README.md CHANGED
@@ -79,6 +79,26 @@ Held-out test split, identical evaluation code for base and fine-tune (base meas
79
 
80
  **Qualitative review** (34 coding/terminal/debugging prompts, judged clean-and-correct): of the prompts that produced a final answer, **92% were correct**. The remainder hit reasoning-budget cutoffs rather than wrong answers (23/34 overall with a 2,600-token budget; see guidance above).
81
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
82
  ## Limitations
83
 
84
  - Like other trace-trained reasoning models, it invests heavily in thinking. With tight token budgets it can spend the whole budget reasoning; budget ≥ 2500 tokens or retry at lower temperature if a response comes back empty.
 
79
 
80
  **Qualitative review** (34 coding/terminal/debugging prompts, judged clean-and-correct): of the prompts that produced a final answer, **92% were correct**. The remainder hit reasoning-budget cutoffs rather than wrong answers (23/34 overall with a 2,600-token budget; see guidance above).
81
 
82
+
83
+ ### Function calling (BFCL V3, AST subset)
84
+
85
+ Measured 2026-07-29: bfcl-eval at gorilla main, prompting mode, Q4_K_M
86
+ GGUFs served by llama.cpp on a T4, base and Parable under the identical
87
+ harness. Categories: simple_python / multiple / parallel /
88
+ parallel_multiple (400/200/200/200 items). Raw generations and score
89
+ files: [parable-v2-artifacts](https://huggingface.co/AnkitAI/parable-v2-artifacts)
90
+ under `verify/bfcl/`.
91
+
92
+ | | simple_python | multiple | parallel | parallel_multiple |
93
+ |---|---|---|---|---|
94
+ | Qwen3-4B base | 0.953 | 0.945 | 0.915 | 0.890 |
95
+ | **Parable-Qwen3-4B** | 0.923 | 0.900 | 0.865 | 0.835 |
96
+
97
+ A 2.3 to 5.5 point trade per category: prose-trace SFT costs a little
98
+ function-calling sharpness, as this card's evaluation note predicts. If
99
+ you need maximum tool-calling accuracy, use the base; this variant buys
100
+ the reasoning voice.
101
+
102
  ## Limitations
103
 
104
  - Like other trace-trained reasoning models, it invests heavily in thinking. With tight token budgets it can spend the whole budget reasoning; budget ≥ 2500 tokens or retry at lower temperature if a response comes back empty.