ChibuUkachi commited on
Commit
c11656c
·
verified ·
1 Parent(s): e7759a5

add BFCL results

Browse files
Files changed (1) hide show
  1. README.md +41 -1
README.md CHANGED
@@ -154,7 +154,10 @@ This model was evaluated on GSM8K-Platinum, MMLU-Pro, IFEval, Math 500, GPQA Dia
154
  | Math 500 | 84.80 | 85.00 | 100.24 |
155
  | Lcb Codegeneration V6 | 77.33 | 74.67 | 96.55 |
156
  | MMLU Pro Chat | 85.32 | 84.70 | 99.28 |
157
-
 
 
 
158
 
159
 
160
  ### Reproduction
@@ -246,5 +249,42 @@ lighteval endpoint litellm litellm_config.yaml \
246
  --save-details
247
  ```
248
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
249
  </details>
250
 
 
154
  | Math 500 | 84.80 | 85.00 | 100.24 |
155
  | Lcb Codegeneration V6 | 77.33 | 74.67 | 96.55 |
156
  | MMLU Pro Chat | 85.32 | 84.70 | 99.28 |
157
+ | BFCLv4 Overall | 57.83 | 56.10 | 97.01% |
158
+ | BFCLv4 Single Turn | 53.81 | 53.45 | 99.34% |
159
+ | BFCLv4 Multi-Turn |62.25 | 58.13 |93.38% |
160
+ | BFCLv4 Agentic |49.91 | 49.31 | 98.80% |
161
 
162
 
163
  ### Reproduction
 
249
  --save-details
250
  ```
251
 
252
+
253
+ #### BFCLv4
254
+
255
+ BFCL requires the model to be registered in the leaderboard codebase before running evaluation.
256
+
257
+ **Step 1 — Register the model in `bfcl_eval/constants/model_config.py`**
258
+
259
+ Add the following entry to `api_inference_model_map`:
260
+
261
+ ```python
262
+ "Qwen3.6-35B-A3B-NVFP4": ModelConfig(
263
+ model_name="Qwen3.6-35B-A3B-NVFP4",
264
+ display_name="Qwen3.6-35B-A3B-NVFP4 (FC)",
265
+ url="https://huggingface.co/RedHatAI/Qwen3.6-35B-A3B-NVFP4",
266
+ org="Google",
267
+ license="Apache 2.0",
268
+ model_handler=OpenAICompletionsHandler,
269
+ input_price=None,
270
+ output_price=None,
271
+ is_fc_model=True,
272
+ underscore_to_dot=True,
273
+ ),
274
+ ```
275
+
276
+ **Step 2 — Add the key to `bfcl_eval/constants/supported_models.py`**
277
+
278
+ Add `"Qwen3.6-35B-A3B-NVFP4"` to the `SUPPORTED_MODELS` list.
279
+
280
+ **Step 3 — Start the vLLM server** (use the command at the top of this section; the `--served-model-name` flag ensures BFCL can find the model by its registered slug).
281
+
282
+ **Step 4 — Generate responses and evaluate**
283
+ ```
284
+ bfcl generate --model Qwen3.6-35B-A3B-NVFP4 --test-category all
285
+ bfcl evaluate --model Qwen3.6-35B-A3B-NVFP4 --test-category all
286
+ ```
287
+
288
+
289
  </details>
290