Frosty40's picture
docs: control A/B ledger + claims hygiene (no GGUF change)
d6cfe8a verified
|
Raw History Blame Contribute Delete
2.98 kB

Benchmarks

Model package: https://huggingface.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF

GitHub: https://github.com/newjordan/treebeard

Site: https://newjordan.github.io/treebeard/

Public result index: https://github.com/newjordan/treebeard/tree/main/results

Agent Bench freeze (named run)

Score 94/100 (130/138 points):

  • 69/69 scenarios completed;
  • 63 pass, 4 partial, 2 fail;
  • zero request errors;
  • one server slot and one benchmark worker;
  • 262,144 total context tokens;
  • temperature 0, thinking disabled, seed 42;
  • tool-eval-bench 2.1.0 at 8b3259b;
  • llama.cpp build b9624-0424f677f (freeze pin).

Same score and outcome vector on Intel Arc Pro B70 and NVIDIA GB10.

Packaged evidence:

  • evidence/agent/single-slot-94/result.json
  • evidence/agent/single-slot-94/index.html
  • evidence/agent/single-slot-94/guard.log
  • evidence/agent/single-slot-94/nvidia-result.json
  • evidence/nvidia/agent-single-slot-94/result.json

Control A/B (2026-07-28, stock Q5, Arc Pro B70)

Same stock Q5 weights, np=1, c=262144, temp 0, seed 42, tool-eval-bench 2.1.0 @ 8b3259b. Clean upstream SYCL versus Treebeard package binary + package env.

Arm Public 69 Held-out ho-pack-v1.1
Original Qwen control 91/100 (125/138) 42/46 (91.3%)
Treebeard package 91/100 (126/138) 42/46 (91.3%)

Sequential tg_p50 (np=1, c=32768, 5 prompts × 2): control 77.1 → Treebeard 88.9. Multi-slot p50/agent (np=12, n_predict=96): control 6.88 → Treebeard 26.33 (concurrent capacity).

Evidence (GitHub, not duplicated as GGUF LFS here):

Portable CPU package smoke

Independent install on AMD Ryzen 9 5950X:

  • build identity: b9624-6a6dc2def-cpu;
  • context: 4,096 tokens for the bounded smoke;
  • chat output: exact TREEBEARD READY;
  • chat generation: 9.302 tok/s;
  • tool call: exact multiply({"a":17,"b":23});
  • tool-call generation: 7.400 tok/s;
  • loader warnings, request errors, and assertion failures: zero.

Functional fallback only. Raw files under evidence/cpu-linux-x86_64.

SYCL serving

Released 12-slot profile: 194.023 aggregate tok/s. 8-slot: 182.005 aggregate tok/s. Single-session aggregate flat at -0.121% vs the prior RC2 reference in the package notes. Supporting evidence under evidence/sycl.

NVIDIA

GB10, compute capability 12.1, CUDA 13.3, Linux ARM64:

  • CUDA correctness: 1,104/1,104 MUL_MAT and 796/796 MUL_MAT_ID;
  • Q8_0 direct 12-column latency: 5.22 to 3.97 us median (31.49% lower);
  • Q8_0 MoE down latency: 463.06 to 445.20 us median (4.01% lower);
  • native pp4096: 2,422.325 tok/s over five samples;
  • native tg128: 59.614 tok/s over five samples;
  • single-slot agent freeze: 94/100, 130/138, zero request errors.

See docs/NVIDIA.md and evidence/nvidia.