# Benchmarks Model package: GitHub: Site: Public result index: ## Agent Bench freeze (named run) Score **94/100** (130/138 points): - 69/69 scenarios completed; - 63 pass, 4 partial, 2 fail; - zero request errors; - one server slot and one benchmark worker; - 262,144 total context tokens; - temperature 0, thinking disabled, seed 42; - tool-eval-bench 2.1.0 at `8b3259b`; - llama.cpp build `b9624-0424f677f` (freeze pin). Same score and outcome vector on Intel Arc Pro B70 and NVIDIA GB10. Packaged evidence: - `evidence/agent/single-slot-94/result.json` - `evidence/agent/single-slot-94/index.html` - `evidence/agent/single-slot-94/guard.log` - `evidence/agent/single-slot-94/nvidia-result.json` - `evidence/nvidia/agent-single-slot-94/result.json` ## Control A/B (2026-07-28, stock Q5, Arc Pro B70) Same stock Q5 weights, np=1, c=262144, temp 0, seed 42, tool-eval-bench 2.1.0 @ `8b3259b`. Clean upstream SYCL versus Treebeard package binary + package env. | Arm | Public 69 | Held-out ho-pack-v1.1 | | --- | ---: | ---: | | Original Qwen control | 91/100 (125/138) | 42/46 (91.3%) | | Treebeard package | 91/100 (126/138) | 42/46 (91.3%) | Sequential tg_p50 (np=1, c=32768, 5 prompts × 2): control 77.1 → Treebeard 88.9. Multi-slot p50/agent (np=12, n_predict=96): control 6.88 → Treebeard 26.33 (concurrent capacity). Evidence (GitHub, not duplicated as GGUF LFS here): - https://github.com/newjordan/treebeard/tree/main/results/private-verification-20260728/agent-bench-ab-20260728T220230Z - https://github.com/newjordan/treebeard/blob/main/docs/RELEASE-20260728.md ## Portable CPU package smoke Independent install on AMD Ryzen 9 5950X: - build identity: `b9624-6a6dc2def-cpu`; - context: 4,096 tokens for the bounded smoke; - chat output: exact `TREEBEARD READY`; - chat generation: 9.302 tok/s; - tool call: exact `multiply({"a":17,"b":23})`; - tool-call generation: 7.400 tok/s; - loader warnings, request errors, and assertion failures: zero. Functional fallback only. Raw files under `evidence/cpu-linux-x86_64`. ## SYCL serving Released 12-slot profile: 194.023 aggregate tok/s. 8-slot: 182.005 aggregate tok/s. Single-session aggregate flat at -0.121% vs the prior RC2 reference in the package notes. Supporting evidence under `evidence/sycl`. ## NVIDIA GB10, compute capability 12.1, CUDA 13.3, Linux ARM64: - CUDA correctness: 1,104/1,104 `MUL_MAT` and 796/796 `MUL_MAT_ID`; - Q8_0 direct 12-column latency: 5.22 to 3.97 us median (31.49% lower); - Q8_0 MoE down latency: 463.06 to 445.20 us median (4.01% lower); - native pp4096: 2,422.325 tok/s over five samples; - native tg128: 59.614 tok/s over five samples; - single-slot agent freeze: 94/100, 130/138, zero request errors. See `docs/NVIDIA.md` and `evidence/nvidia`.