Benchmarked this quant on BFCL v4: 89% non-live function-calling, fits 16GB with 65k context

#42
by tumbak - opened

Thanks for shipping this one β€” the 2026-05-04 UD-IQ4_XS with the official Gemma chat template baked in is what I ended up running daily.

I put it through BFCL v4 (Berkeley Function Calling Leaderboard) to get real agentic tool-use numbers rather than vibes. As far as I can tell these are the first published BFCL figures for Gemma 4. Sharing here since they're specific to this quant.

Setup: single RTX 5070 Ti (16GB), llama.cpp + --jinja, prompt mode, temperature 1.0 (Google's recommended sampling), UD-IQ4_XS (12.65 GiB).

BFCL v4 (overall) Score
Non-Live 89.13%
Live 63.80%
Multi-Turn 45.12%

For reference, simple_python lands at 95% β€” same band as models several times its active parameter count. The 12.65 GiB footprint leaves room for a 65k context window on 16GB, which is the part that makes it actually usable locally.

One thing worth flagging for anyone using this GGUF for agentic/multi-turn work: under --jinja, Gemma 4 emits tool calls in its native syntax (<|tool_call>call:fn(args)<tool_call|>), and its chat template silently drops role="tool" messages β€” so naive multi-turn harnesses never feed tool results back and the model loops blind. Both are fixable on the harness side.

Full methodology, throughput numbers, the build saga, and per-category scores: https://algollabs.com/blog/gemma4-bfcl

I also upstreamed a Gemma 4 handler to BFCL so others can reproduce: https://github.com/ShishirPatil/gorilla/pull/1340

Question for the thread: has anyone run vanilla IQ4_XS vs this dynamic UD-IQ4_XS head-to-head on a structured benchmark? My April (vanilla) vs May (UD) deltas were within single-seed noise, but quant and chat template changed at the same time so I couldn't cleanly attribute. Curious if anyone has a clean A/B.

Sign up or log in to comment