gowitheflow's picture
Add files using upload-large-folder tool
f5568aa verified
|
Raw History Blame
1.96 kB
---
pretty_name: AutoDataBench Function Calling Resources
tags:
- autodatabench
- function-calling
- tool-use
---
# AutoDataBench Function Calling Resources
Public resources for the function-calling task in
[AutoDataBench](https://github.com/AutoDataBench/AutoDataBench). See the
[paper](https://arxiv.org/abs/2609.40097) for the benchmark setting.
## Contents
```text
data/function_call_v1/pool_agent.jsonl
models/Qwen2-1.5B-Instruct/
models/Qwen3-4B-Instruct-2507/
models/Qwen3-Embedding-0.6B/
```
`pool_agent.jsonl` is the 40,001-row noisy single-turn function-calling pool
available to the data agent. It does not expose the noise labels used to build
the pool.
| Model | Role | Original model |
| --- | --- | --- |
| Qwen2-1.5B-Instruct | Fixed function-calling base model | [Qwen/Qwen2-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2-1.5B-Instruct) |
| Qwen3-4B-Instruct-2507 | Agent-callable generation model | [Qwen/Qwen3-4B-Instruct-2507](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507) |
| Qwen3-Embedding-0.6B | Agent-callable embedding model | [Qwen/Qwen3-Embedding-0.6B](https://huggingface.co/Qwen/Qwen3-Embedding-0.6B) |
## Evaluation data
The held-out in-domain test set is intentionally excluded. The BFCL v3 OOD
guard is also not duplicated here; maintainers should obtain it from the
[Berkeley Function Calling Leaderboard](https://huggingface.co/datasets/gorilla-llm/Berkeley-Function-Calling-Leaderboard)
and keep gold answers outside the agent sandbox.
## Use with AutoDataBench
Copy or symlink `data/` and `models/` into the AutoDataBench repository. The
paths already match the default task configuration. Point the generation and
embedding servers at the local auxiliary-model directories if needed.
`MANIFEST.sha256` contains checksums for every distributed file.
Model and dataset components retain their upstream licenses. Consult the model
cards and source datasets before redistribution or commercial use.