# Tool-Use Benchmark for Small/Efficient LLMs on Apple Silicon Benchmarking tool-use (function calling) capability of small and efficient language models running locally on a Mac Mini M4 (16 GB). Covers 13 model configurations across 3 backends, evaluated on 3 benchmarks. ## Key Finding **PrismML Bonsai-8B (1-bit, 1.15 GB) achieves 73.3% on BFCL** — beating Qwen3.5-9B (61%), Gemma 4 E4B (65%), and every other model tested. A 1-bit quantized model at 14x compression outperforms 4-bit models 4-5x its size on structured function calling. However, on complex API semantics (NexusRaven), Bonsai-8B drops to 43.8% vs Qwen3.5-9B's 75%. **BFCL measures format compliance; NexusRaven measures understanding.** ## Results ### BFCL (Berkeley Function Calling Leaderboard) — 50 tests per category | Model | Size | Quant | Backend | Simple | Multiple | Parallel | **Avg** | Avg Time | |---|---|---|---|---|---|---|---|---| | **Bonsai-8B** | 1.15 GB | Q1_0 1-bit | llama.cpp | 68% | 72% | 80% | **73.3%** | 1.8s | | Gemma 4 E4B-it | ~5 GB | Q4_K_M | Ollama | 54% | 64% | 78% | **65.3%** | 2.4s | | Qwen3.5-9B | ~5 GB | Q4_K_M | llama.cpp | 56% | 68% | 68% | **64.0%** | 11.6s | | Qwen3.5-9B | ~5 GB | MLX 4-bit | mlx-vlm | 60% | 68% | 64% | **64.0%** | 9.5s | | Qwen2.5-7B | ~4.7 GB | Q4_K_M | Ollama | 58% | 62% | 70% | **63.3%** | 2.9s | | Gemma 4 E2B-it | ~3 GB | Q4_K_M | Ollama | 56% | 60% | 70% | **62.0%** | 1.3s | | Gemma 3 12B | ~7.3 GB | Q4_K_M | Ollama | 54% | 54% | 78% | **62.0%** | 5.4s | | Qwen3.5-9B | ~5 GB | Q4_K_M | Ollama | 50% | 60% | 74% | **61.3%** | 5.4s | | Bonsai-4B | 0.57 GB | Q1_0 1-bit | llama.cpp | 36% | 56% | 74% | **55.3%** | 1.0s | | Bonsai-1.7B | 0.25 GB | Q1_0 1-bit | llama.cpp | 58% | 54% | 54% | **55.3%** | 0.4s | | Llama 3.1 8B | ~4.7 GB | Q4_K_M | Ollama | 46% | 42% | 66% | **51.3%** | 3.0s | | Mistral-Nemo 12B | ~7.1 GB | Q4_K_M | Ollama | 40% | 44% | 64% | **49.3%** | 4.4s | | Bonsai-4B FP16 | 7.5 GB | FP16 | mlx-lm | 8% | 34% | 34% | **25.3%** | 4.8s | ### NexusRaven API Evaluation — 48 queries, 12 per domain, stratified | Model | Size | Overall | cve_cpe | emailrep | virustotal | toolalpaca | Avg Time | |---|---|---|---|---|---|---|---| | Qwen3.5-9B (llama.cpp) | ~5 GB | **77.1%** | 58% | 100% | 100% | 50% | 14.1s | | Qwen3.5-9B (Ollama) | ~5 GB | **75.0%** | 58% | 100% | 100% | 42% | 4.1s | | Qwen2.5-7B | ~4.7 GB | **70.8%** | 50% | 92% | 100% | 42% | 2.0s | | Qwen3.5-9B (mlx-vlm) | ~5 GB | **70.8%** | 50% | 100% | 92% | 42% | 13.8s | | Gemma 3 12B | ~7.3 GB | **68.8%** | 33% | 100% | 100% | 42% | 3.5s | | Llama 3.1 8B | ~4.7 GB | **66.7%** | 25% | 92% | 100% | 50% | 2.1s | | Mistral-Nemo 12B | ~7.1 GB | **66.7%** | 42% | 92% | 100% | 33% | 3.0s | | Gemma 4 E4B-it | ~5 GB | **60.4%** | 33% | 83% | 83% | 42% | 1.6s | | Bonsai-1.7B | 0.25 GB | **54.2%** | 25% | 75% | 83% | 33% | 0.3s | | Gemma 4 E2B-it | ~3 GB | **47.9%** | 17% | 67% | 75% | 33% | 0.9s | | Bonsai-4B | 0.57 GB | **43.8%** | 17% | 58% | 75% | 25% | 0.8s | | Bonsai-8B | 1.15 GB | **43.8%** | 17% | 67% | 67% | 25% | 1.2s | | Bonsai-4B FP16 | 7.5 GB | **29.2%** | 8% | 42% | 50% | 17% | 3.5s | ### AgentBench OS — Multi-step agentic tasks in Docker containers | Model | Score | Backend | |---|---|---| | **Qwen3.5-9B** | **4.5/10 (45%)** | Ollama / llama.cpp (tied) | | Qwen2.5-7B | 3.5/10 (35%) | Ollama | | Mistral-Nemo 12B | 3.5/10 (35%) | Ollama | | Llama 3.1 8B | 3.0/10 (30%) | Ollama | | Gemma 3 12B | 2.5/10 (25%) | Ollama | ### Backend Comparison (Qwen3.5-9B only, all 3 benchmarks) | Backend | AgentBench | BFCL | NexusRaven | Composite | |---|---|---|---|---| | llama.cpp (UD-Q4_K_XL) | 45% | 64.0% | 77.1% | **62.0%** | | Ollama (Q4_K_M) | 45% | 61.3% | 75.0% | **60.4%** | | mlx-vlm (MLX-4bit) | 42% | 64.0% | 70.8% | **58.9%** | ## Hardware - **Mac Mini M4** (10-core CPU, 10-core GPU, 16 GB unified memory) - macOS Sequoia 15.3 - All models run locally, no cloud APIs ## Benchmarks ### 1. BFCL (Berkeley Function Calling Leaderboard) - **What**: Structured function calling — can the model output the correct function name and parameters? - **Dataset**: [gorilla-llm/Berkeley-Function-Calling-Leaderboard](https://huggingface.co/datasets/gorilla-llm/Berkeley-Function-Calling-Leaderboard) v3 - **Categories**: `simple` (1 function), `multiple` (choose from many), `parallel` (multiple calls at once) - **Metric**: Exact match on function name + type-flexible parameter matching ### 2. NexusRaven API Evaluation - **What**: Real-world API calling with complex parameter schemas (up to 28 params) - **Dataset**: [Nexusflow/NexusRaven_API_evaluation](https://huggingface.co/datasets/Nexusflow/NexusRaven_API_evaluation) - **Domains**: CVE/CPE (security), EmailRep, VirusTotal, ToolAlpaca - **Metric**: Correct function + all required parameters with correct values ### 3. AgentBench OS - **What**: Multi-step agentic tasks in Docker containers (file manipulation, data processing, system admin) - **Dataset**: Adapted from [THUDM/AgentBench](https://github.com/THUDM/AgentBench) OS interaction subset - **Metric**: Task completion scored by automated verifiers ## Models Tested | Model | Parameters | Format | Size on Disk | |---|---|---|---| | [PrismML Bonsai-8B](https://huggingface.co/prism-ml/Bonsai-8B-gguf) | 8B | Q1_0 1-bit GGUF | 1.15 GB | | [PrismML Bonsai-4B](https://huggingface.co/prism-ml/Bonsai-4B-gguf) | 4B | Q1_0 1-bit GGUF | 0.57 GB | | [PrismML Bonsai-1.7B](https://huggingface.co/prism-ml/Bonsai-1.7B-gguf) | 1.7B | Q1_0 1-bit GGUF | 0.25 GB | | Bonsai-4B FP16 | 4B | FP16 Safetensors | 7.5 GB | | [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3-8B) | 9B | Q4_K_M / MLX-4bit / UD-Q4_K_XL | ~5 GB | | [Qwen2.5-7B](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) | 7B | Q4_K_M | ~4.7 GB | | [Gemma 4 E4B-it](https://huggingface.co/google/gemma-3-4b-it) | 4B | Q4_K_M | ~5 GB | | [Gemma 4 E2B-it](https://huggingface.co/google/gemma-3-2b-it) | 2B | Q4_K_M | ~3 GB | | [Gemma 3 12B](https://huggingface.co/google/gemma-3-12b-it) | 12B | Q4_K_M | ~7.3 GB | | [Llama 3.1 8B](https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct) | 8B | Q4_K_M | ~4.7 GB | | [Mistral-Nemo 12B](https://huggingface.co/mistralai/Mistral-Nemo-Instruct-2407) | 12B | Q4_K_M | ~7.1 GB | ## Backends | Backend | Version | Notes | |---|---|---| | [Ollama](https://ollama.com) | v0.20+ | OpenAI-compatible at `/api/chat` | | [llama.cpp](https://github.com/ggml-org/llama.cpp) | b8640 | Via `llama-server`, OpenAI-compatible at `/v1/chat/completions` | | [PrismML llama.cpp](https://github.com/PrismML-Eng/llama.cpp) | fork | Required for Bonsai Q1_0 (ggml type 41) | | [mlx-vlm](https://github.com/Blaizzy/mlx-vlm) | latest | MLX-native, API at `/chat/completions` | ## Reproduction ### Prerequisites - macOS with Apple Silicon (M1+, tested on M4) - Python 3.10+ - `pip install requests datasets` - One or more backends installed (Ollama, llama.cpp, mlx-vlm) ### Run BFCL ```bash # First run downloads the dataset from HuggingFace (~200 MB) # Ollama backend python3 scripts/run_bfcl.py --model qwen3.5:9b --backend ollama \ --categories simple,multiple,parallel --limit 50 # llama.cpp backend (start server first) llama-server -m model.gguf --port 8081 -ngl 99 -c 4096 python3 scripts/run_bfcl.py --model bonsai-8b --backend llama-cpp \ --categories simple,multiple,parallel --limit 50 ``` ### Run NexusRaven ```bash # Downloads NexusRaven dataset from HuggingFace on first run python3 scripts/run_nexusraven.py --model qwen3.5:9b --backend ollama --limit 48 python3 scripts/run_nexusraven.py --model bonsai-8b --backend llama-cpp --limit 48 ``` ### Run AgentBench OS ```bash # Requires Docker (via Colima on macOS) brew install colima docker colima start --memory 4 python3 scripts/run_bench.py --backend ollama --model qwen3.5:9b \ --dataset datasets/agentbench_os_v1.json ``` ### Bonsai Models (PrismML fork required) Stock llama.cpp does not support Q1_0_g128 (ggml type 41). Build the PrismML fork: ```bash git clone https://github.com/PrismML-Eng/llama.cpp prism-llama-cpp cd prism-llama-cpp cmake -B build && cmake --build build -j # Start server ./build/bin/llama-server -m Bonsai-8B.gguf --port 8081 -ngl 99 -c 4096 ``` ## Repo Structure ``` tool-use-bench/ README.md scripts/ run_bfcl.py # BFCL evaluator (multi-backend) run_nexusraven.py # NexusRaven evaluator (multi-backend) run_bench.py # AgentBench OS evaluator (Docker) results/ bfcl/ # 13 result JSONs (per model+backend) nexusraven/ # 13 result JSONs agentbench/ # Per-backend and per-model results ollama/ llama-cpp/ mlx-vlm/ qwen2.5_7b/ gemma3_12b/ llama3.1_8b/ mistral-nemo_12b/ datasets/ agentbench_os_v1.json # 10 OS interaction tasks (set 1) agentbench_os_v2.json # 10 OS interaction tasks (set 2) agentbench_os_v3.json # 10 OS interaction tasks (set 3) agentbench_os_v4.json # 10 OS interaction tasks (set 4) ``` ## Insights 1. **1-bit quantization can preserve tool-use capability** — Bonsai's quantization-aware training actually *improves* structured output (73% BFCL) compared to its FP16 base (25%). The 1-bit model is both smaller and better. 2. **BFCL and NexusRaven measure different things** — Bonsai excels at format compliance (BFCL) but struggles with complex API semantics (NexusRaven). For edge deployment where API schemas are fixed and simple, Bonsai is excellent. For dynamic/complex APIs, Qwen3.5-9B is the better choice. 3. **Backend choice barely matters** — Ollama, llama.cpp, and mlx-vlm produce nearly identical accuracy (within ~3%). The model is the bottleneck, not the serving infrastructure. 4. **Size vs capability trade-off** — Bonsai-1.7B (0.25 GB, 0.4s/query) scores 55% on BFCL and 54% on NexusRaven. For the price of a single image, you get a competent function-calling model that runs on a phone. ## License Scripts are MIT licensed. Benchmark datasets are subject to their original licenses: - BFCL: Apache 2.0 ([gorilla-llm](https://huggingface.co/datasets/gorilla-llm/Berkeley-Function-Calling-Leaderboard)) - NexusRaven: Apache 2.0 ([Nexusflow](https://huggingface.co/datasets/Nexusflow/NexusRaven_API_evaluation)) - AgentBench: MIT ([THUDM](https://github.com/THUDM/AgentBench)) ## Citation If you use these results or scripts, please cite: ``` @misc{tool-use-bench-m4, title={Tool-Use Benchmark for Small LLMs on Apple Silicon}, year={2026}, url={https://github.com/user/tool-use-bench} } ```