File size: 2,977 Bytes
247b1b9
 
d6cfe8a
247b1b9
d6cfe8a
247b1b9
d6cfe8a
247b1b9
d6cfe8a
 
 
 
 
247b1b9
 
 
 
 
 
 
 
d6cfe8a
 
 
247b1b9
d6cfe8a
247b1b9
 
 
 
 
 
 
d6cfe8a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
247b1b9
 
 
d6cfe8a
247b1b9
 
 
 
 
 
 
 
 
d6cfe8a
247b1b9
 
 
d6cfe8a
 
 
247b1b9
 
 
d6cfe8a
247b1b9
 
d6cfe8a
 
247b1b9
 
d6cfe8a
247b1b9
d6cfe8a
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
# Benchmarks

Model package: <https://huggingface.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF>

GitHub: <https://github.com/newjordan/treebeard>

Site: <https://newjordan.github.io/treebeard/>

Public result index: <https://github.com/newjordan/treebeard/tree/main/results>

## Agent Bench freeze (named run)

Score **94/100** (130/138 points):

- 69/69 scenarios completed;
- 63 pass, 4 partial, 2 fail;
- zero request errors;
- one server slot and one benchmark worker;
- 262,144 total context tokens;
- temperature 0, thinking disabled, seed 42;
- tool-eval-bench 2.1.0 at `8b3259b`;
- llama.cpp build `b9624-0424f677f` (freeze pin).

Same score and outcome vector on Intel Arc Pro B70 and NVIDIA GB10.

Packaged evidence:

- `evidence/agent/single-slot-94/result.json`
- `evidence/agent/single-slot-94/index.html`
- `evidence/agent/single-slot-94/guard.log`
- `evidence/agent/single-slot-94/nvidia-result.json`
- `evidence/nvidia/agent-single-slot-94/result.json`

## Control A/B (2026-07-28, stock Q5, Arc Pro B70)

Same stock Q5 weights, np=1, c=262144, temp 0, seed 42, tool-eval-bench 2.1.0
@ `8b3259b`. Clean upstream SYCL versus Treebeard package binary + package env.

| Arm | Public 69 | Held-out ho-pack-v1.1 |
| --- | ---: | ---: |
| Original Qwen control | 91/100 (125/138) | 42/46 (91.3%) |
| Treebeard package | 91/100 (126/138) | 42/46 (91.3%) |

Sequential tg_p50 (np=1, c=32768, 5 prompts × 2): control 77.1 → Treebeard 88.9.
Multi-slot p50/agent (np=12, n_predict=96): control 6.88 → Treebeard 26.33
(concurrent capacity).

Evidence (GitHub, not duplicated as GGUF LFS here):

- https://github.com/newjordan/treebeard/tree/main/results/private-verification-20260728/agent-bench-ab-20260728T220230Z
- https://github.com/newjordan/treebeard/blob/main/docs/RELEASE-20260728.md

## Portable CPU package smoke

Independent install on AMD Ryzen 9 5950X:

- build identity: `b9624-6a6dc2def-cpu`;
- context: 4,096 tokens for the bounded smoke;
- chat output: exact `TREEBEARD READY`;
- chat generation: 9.302 tok/s;
- tool call: exact `multiply({"a":17,"b":23})`;
- tool-call generation: 7.400 tok/s;
- loader warnings, request errors, and assertion failures: zero.

Functional fallback only. Raw files under `evidence/cpu-linux-x86_64`.

## SYCL serving

Released 12-slot profile: 194.023 aggregate tok/s. 8-slot: 182.005 aggregate
tok/s. Single-session aggregate flat at -0.121% vs the prior RC2 reference in
the package notes. Supporting evidence under `evidence/sycl`.

## NVIDIA

GB10, compute capability 12.1, CUDA 13.3, Linux ARM64:

- CUDA correctness: 1,104/1,104 `MUL_MAT` and 796/796 `MUL_MAT_ID`;
- Q8_0 direct 12-column latency: 5.22 to 3.97 us median (31.49% lower);
- Q8_0 MoE down latency: 463.06 to 445.20 us median (4.01% lower);
- native pp4096: 2,422.325 tok/s over five samples;
- native tg128: 59.614 tok/s over five samples;
- single-slot agent freeze: 94/100, 130/138, zero request errors.

See `docs/NVIDIA.md` and `evidence/nvidia`.