Frosty40 commited on
Commit
98dd39a
·
1 Parent(s): eebac6c

Treebeard 0.1.1 launcher refresh (pkg4): reasoning + speculation controls

Browse files

run.sh/treebeard pkg4: TREEBEARD_REASONING=off|bounded|unrestricted and
TREEBEARD_SPECULATION=off|ngram|mtp|hybrid; PACKAGE.json optional_controls;
card + PACKAGE.md docs; SHA256SUMS re-pinned. Runtime bytes unchanged.

Files changed (6) hide show
  1. PACKAGE.json +26 -1
  2. README.md +14 -1
  3. SHA256SUMS +5 -5
  4. docs/PACKAGE.md +43 -0
  5. run.sh +89 -3
  6. treebeard +3 -0
PACKAGE.json CHANGED
@@ -3,7 +3,7 @@
3
  "name": "Treebeard-Qwen3.6-35B-A3B-GGUF",
4
  "product": "Treebeard",
5
  "version": "0.1.0-rc.3",
6
- "packaging_revision": "pkg3",
7
  "visibility": "public",
8
  "created_date": "2026-07-11",
9
  "publication": {
@@ -103,6 +103,31 @@
103
  "server_slots": 1
104
  }
105
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
106
  "agent_benchmark": {
107
  "tool": "tool-eval-bench 2.1.0",
108
  "commit": "8b3259be7411fe27c7610d0de64ae1d3b622b9ef",
 
3
  "name": "Treebeard-Qwen3.6-35B-A3B-GGUF",
4
  "product": "Treebeard",
5
  "version": "0.1.0-rc.3",
6
+ "packaging_revision": "pkg4",
7
  "visibility": "public",
8
  "created_date": "2026-07-11",
9
  "publication": {
 
103
  "server_slots": 1
104
  }
105
  },
106
+ "optional_controls": {
107
+ "reasoning": {
108
+ "default": "off",
109
+ "modes": [
110
+ "off",
111
+ "bounded",
112
+ "unrestricted"
113
+ ],
114
+ "bounded_budget_tokens": {
115
+ "gpu": 64,
116
+ "cpu": 16
117
+ }
118
+ },
119
+ "speculation": {
120
+ "default": "off",
121
+ "modes": [
122
+ "off",
123
+ "ngram",
124
+ "mtp",
125
+ "hybrid"
126
+ ],
127
+ "native_mtp_layers": 1,
128
+ "validated_speedup": false
129
+ }
130
+ },
131
  "agent_benchmark": {
132
  "tool": "tool-eval-bench 2.1.0",
133
  "commit": "8b3259be7411fe27c7610d0de64ae1d3b622b9ef",
README.md CHANGED
@@ -21,7 +21,7 @@ pipeline_tag: image-text-to-text
21
  # Treebeard
22
 
23
  Treebeard is a ready-to-run Linux package for Qwen3.6-35B-A3B. Release
24
- `0.1.0-rc.3` combines one 26.6 GB Q5_K_XL GGUF, the optional F16 vision
25
  projector, official Qwen metadata and tokenizer files, platform runtimes, a
26
  one-command installer, an OpenAI-compatible server launcher, and the raw
27
  evidence behind its benchmark claims.
@@ -152,6 +152,9 @@ tokens. Override settings with environment variables:
152
  | `TREEBEARD_HOST` | `127.0.0.1` | API bind address |
153
  | `TREEBEARD_PORT` | `8093` | API port |
154
  | `TREEBEARD_MULTIMODAL` | `0` | Load the installed F16 projector |
 
 
 
155
  | `TREEBEARD_VERIFY` | `once` | `once`, `always`, or `never` |
156
  | `TREEBEARD_CPUSET` | unset | Optional `taskset` CPU list |
157
 
@@ -168,6 +171,16 @@ not the configuration behind the single-slot 94:
168
  TREEBEARD_PROFILE=throughput treebeard serve
169
  ```
170
 
 
 
 
 
 
 
 
 
 
 
171
  ## Full repository package
172
 
173
  The Hugging Face repository is the complete package. After downloading it with
 
21
  # Treebeard
22
 
23
  Treebeard is a ready-to-run Linux package for Qwen3.6-35B-A3B. Release
24
+ `0.1.1` (runtime `0.1.0-rc.3`, launcher pkg4) combines one 26.6 GB Q5_K_XL GGUF, the optional F16 vision
25
  projector, official Qwen metadata and tokenizer files, platform runtimes, a
26
  one-command installer, an OpenAI-compatible server launcher, and the raw
27
  evidence behind its benchmark claims.
 
152
  | `TREEBEARD_HOST` | `127.0.0.1` | API bind address |
153
  | `TREEBEARD_PORT` | `8093` | API port |
154
  | `TREEBEARD_MULTIMODAL` | `0` | Load the installed F16 projector |
155
+ | `TREEBEARD_REASONING` | `off` | `off`, `bounded`, or `unrestricted` |
156
+ | `TREEBEARD_REASONING_BUDGET` | `64` GPU / `16` CPU | Bounded thinking tokens |
157
+ | `TREEBEARD_SPECULATION` | `off` | `off`, `ngram`, `mtp`, or `hybrid` |
158
  | `TREEBEARD_VERIFY` | `once` | `once`, `always`, or `never` |
159
  | `TREEBEARD_CPUSET` | unset | Optional `taskset` CPU list |
160
 
 
171
  TREEBEARD_PROFILE=throughput treebeard serve
172
  ```
173
 
174
+ Reasoning is off by default, matching the validated 94/100 agent benchmark.
175
+ `TREEBEARD_REASONING=bounded` enables a small thinking allowance (64 tokens
176
+ on GPU, 16 on CPU; override with `TREEBEARD_REASONING_BUDGET`), and
177
+ `unrestricted` removes the budget. API clients can still opt individual
178
+ requests into thinking with request-level chat-template and thinking-budget
179
+ controls. Speculative decoding is opt-in through
180
+ `TREEBEARD_SPECULATION=ngram|mtp|hybrid` (`mtp` uses Qwen3.6's native
181
+ one-layer MTP head carried by the GGUF); these modes ship in the 0.1.1
182
+ launcher refresh but are not yet validated as a Treebeard speedup.
183
+
184
  ## Full repository package
185
 
186
  The Hugging Face repository is the complete package. After downloading it with
SHA256SUMS CHANGED
@@ -2,15 +2,15 @@ cae3aa82fc25bb3f9f6ae4c826f8c08dea0bba30658c448a7f124737549ce345 .gitattributes
2
  50cbab8a892c5f2993b8c7351a99182507472def3b1374558308605d99b86b32 LICENSE
3
  94f29bbed6a22c35b992c5c6ebf0e7c92f13b836b90f36f461c9cf2f0f1d010d LICENSE-RUNTIME
4
  2d04af27c51a8ba55b9f0cbfcf1df72270daf70cf5e5f6d6cbd2db794e9c2151 NOTICE.md
5
- 1170075b0132584c68cd0f3d8d625e5c493a837ccb0d75adf4c9d5e70c4bc5de PACKAGE.json
6
  25233af7642e3a91bd52cc4aeefdbd4a117479088e06cf1aea5b6bedb443c506 Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
7
- a2cddd57a1366568d96a3564f86e6f413cc649b7b815d8028033009203a52692 README.md
8
  e84f32a23fdda27689f868aa4a1a5621f41133e51a48d7f3efcbea2839574259 chat_template.jinja
9
  93a4693fa9d8392fbfccd4b3c9873f4bfdcb14fdede978b123d07d19675efe99 config.json
10
  c1b09db419119513247e9b8b912c4b9897106c9b20c6cada7e107d993c5435eb configuration.json
11
  ae1c5fb00a620a316159929e229e4ca6a170efa6317e55164e614b120af417fd docs/BENCHMARKS.md
12
  f648d15c3c111b72f775f9f08a43611d8398a9ff1855aba86d50576096a8b0c8 docs/NVIDIA.md
13
- b2338532cb94c426843d75bb9ad80b46f03bd788f6634da19c536579bb72e00b docs/PACKAGE.md
14
  0daf18022e86b6e12510735aefa8ec52fea1c48b2f5c1136608287d48bda3f99 evidence/README.md
15
  efc91f319aef7a32c311a4d0f058dd87c92c1b96352fc322966c8ce2ed6fd4eb evidence/agent/single-slot-94/LEGACY-RUN-IDENTITY.md
16
  6e3b2f8e317b123ae8fc35d91ca44a3765fdba1a8d50502be59624a6e24d3809 evidence/agent/single-slot-94/SHA256SUMS
@@ -143,14 +143,14 @@ a9d356d7bdf1ef4949e3e748e95b8e10ad9d4e2e838eddc38a0a7b6b94d1db8d merges.txt
143
  27225450ac9c6529872ee1924fcb0962ff5634834f817040f444118116f4e516 preprocessor_config.json
144
  bcbcf46eae0ada57cfa71645004fd0850781ef2bad7c9da57f7b922169baa421 profiles/quality.env
145
  97afe010ba8d3292456f875d115f03a257a95c12ae02d145c2bcaef449f4718e profiles/throughput.env
146
- bae9586444cb35f2303f936cd0fffd47a845219e4bebd501a3133d45ffed1831 run.sh
147
  48f8c4abb9467fcd0e4c0e8284251903f0e40b27dbf20786f3d9969478c0c81a runtime/SHA256SUMS
148
  51497d7a72cfe8cdf3d66fcbf3ba36f8d5130084cc339cdc49995ec131126732 runtime/cpu-linux-x86_64/treebeard-0.1.0-rc.3-b9624-pkg3-cpu-ubuntu22.04-linux-x86_64.tar.gz
149
  0dfadb9ed53c66ab5a23b6731988d833b223af4e51c8f11ad9522a95eb0f04f3 runtime/cuda-linux-aarch64/treebeard-0.1.0-rc.3-b9624-pkg2-cuda13.3-linux-aarch64.tar.gz
150
  beec6bd316cc5285481d1545e66b31f984b1eb93e69fe05c7d0b1c5e1bafc756 runtime/sycl-linux-x86_64/treebeard-0.1.0-rc.3-b9624-pkg2-sycl-oneapi2026-linux-x86_64.tar.gz
151
  5f9e4d4901a92b997e463c1f46055088b6cca5ca61a6522d1b9f64c4bb81cb42 tokenizer.json
152
  5186f0defcd7f232382c7f0aebcd2252d073bb921ab240e407b7ae8745d2b29b tokenizer_config.json
153
- e372e1c0faba070e61d6412607de6e6149cf0725b1c0a238f96aedbc61d5a3e0 treebeard
154
  3660ebd0bf0e9ce8c3a73af312077029793c3d9a8f52d973c7cd5fc9ca5e5f63 verify.sh
155
  7768af27c1fafa9cc9011c1dc20067e03f8915e03b63504550e11d5066986d13 video_preprocessor_config.json
156
  ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003 vocab.json
 
2
  50cbab8a892c5f2993b8c7351a99182507472def3b1374558308605d99b86b32 LICENSE
3
  94f29bbed6a22c35b992c5c6ebf0e7c92f13b836b90f36f461c9cf2f0f1d010d LICENSE-RUNTIME
4
  2d04af27c51a8ba55b9f0cbfcf1df72270daf70cf5e5f6d6cbd2db794e9c2151 NOTICE.md
5
+ 7bdb30865738053188af49629d2fd9c0fbe9843d7c913727df01c7e554214f69 PACKAGE.json
6
  25233af7642e3a91bd52cc4aeefdbd4a117479088e06cf1aea5b6bedb443c506 Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
7
+ a755e93db87862c16b9765b97a3bb9e2617d767a2aaa8565699e24a0257fa9aa README.md
8
  e84f32a23fdda27689f868aa4a1a5621f41133e51a48d7f3efcbea2839574259 chat_template.jinja
9
  93a4693fa9d8392fbfccd4b3c9873f4bfdcb14fdede978b123d07d19675efe99 config.json
10
  c1b09db419119513247e9b8b912c4b9897106c9b20c6cada7e107d993c5435eb configuration.json
11
  ae1c5fb00a620a316159929e229e4ca6a170efa6317e55164e614b120af417fd docs/BENCHMARKS.md
12
  f648d15c3c111b72f775f9f08a43611d8398a9ff1855aba86d50576096a8b0c8 docs/NVIDIA.md
13
+ 999acc5f1949ed80f5775503821b7f81ab1c31b47581e0ab884139dc49435398 docs/PACKAGE.md
14
  0daf18022e86b6e12510735aefa8ec52fea1c48b2f5c1136608287d48bda3f99 evidence/README.md
15
  efc91f319aef7a32c311a4d0f058dd87c92c1b96352fc322966c8ce2ed6fd4eb evidence/agent/single-slot-94/LEGACY-RUN-IDENTITY.md
16
  6e3b2f8e317b123ae8fc35d91ca44a3765fdba1a8d50502be59624a6e24d3809 evidence/agent/single-slot-94/SHA256SUMS
 
143
  27225450ac9c6529872ee1924fcb0962ff5634834f817040f444118116f4e516 preprocessor_config.json
144
  bcbcf46eae0ada57cfa71645004fd0850781ef2bad7c9da57f7b922169baa421 profiles/quality.env
145
  97afe010ba8d3292456f875d115f03a257a95c12ae02d145c2bcaef449f4718e profiles/throughput.env
146
+ acd5c33ea206ba86814c5c739647a08f80b693332354c7a753a1eac9bade9206 run.sh
147
  48f8c4abb9467fcd0e4c0e8284251903f0e40b27dbf20786f3d9969478c0c81a runtime/SHA256SUMS
148
  51497d7a72cfe8cdf3d66fcbf3ba36f8d5130084cc339cdc49995ec131126732 runtime/cpu-linux-x86_64/treebeard-0.1.0-rc.3-b9624-pkg3-cpu-ubuntu22.04-linux-x86_64.tar.gz
149
  0dfadb9ed53c66ab5a23b6731988d833b223af4e51c8f11ad9522a95eb0f04f3 runtime/cuda-linux-aarch64/treebeard-0.1.0-rc.3-b9624-pkg2-cuda13.3-linux-aarch64.tar.gz
150
  beec6bd316cc5285481d1545e66b31f984b1eb93e69fe05c7d0b1c5e1bafc756 runtime/sycl-linux-x86_64/treebeard-0.1.0-rc.3-b9624-pkg2-sycl-oneapi2026-linux-x86_64.tar.gz
151
  5f9e4d4901a92b997e463c1f46055088b6cca5ca61a6522d1b9f64c4bb81cb42 tokenizer.json
152
  5186f0defcd7f232382c7f0aebcd2252d073bb921ab240e407b7ae8745d2b29b tokenizer_config.json
153
+ 16eb404820fb872f14460a546e79f3cf99c98d2314b40167c105ed08f2f1c132 treebeard
154
  3660ebd0bf0e9ce8c3a73af312077029793c3d9a8f52d973c7cd5fc9ca5e5f63 verify.sh
155
  7768af27c1fafa9cc9011c1dc20067e03f8915e03b63504550e11d5066986d13 video_preprocessor_config.json
156
  ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003 vocab.json
docs/PACKAGE.md CHANGED
@@ -1,5 +1,11 @@
1
  # Package contract
2
 
 
 
 
 
 
 
3
  Treebeard pkg3 follows a standard Hugging Face GGUF layout. The complete GGUF
4
  model and official Qwen configuration and tokenizer files live at repository
5
  root. Runtimes, launch tools, and evidence are additive directories.
@@ -88,3 +94,40 @@ correctness where applicable, and a real server/API package smoke test.
88
  The model supports up to 262,144 context tokens, but actual usable context is
89
  bounded by available device or system memory. The launcher exposes overrides
90
  instead of claiming every host can sustain the maximum profile.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  # Package contract
2
 
3
+ Model package: <https://huggingface.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF>
4
+
5
+ GitHub repository: <https://github.com/newjordan/treebeard>
6
+
7
+ MoE algorithm explainer: <https://newjordan.github.io/treebeard/moe-routing.html>
8
+
9
  Treebeard pkg3 follows a standard Hugging Face GGUF layout. The complete GGUF
10
  model and official Qwen configuration and tokenizer files live at repository
11
  root. Runtimes, launch tools, and evidence are additive directories.
 
94
  The model supports up to 262,144 context tokens, but actual usable context is
95
  bounded by available device or system memory. The launcher exposes overrides
96
  instead of claiming every host can sustain the maximum profile.
97
+
98
+ ## Reasoning and speculative decoding
99
+
100
+ The validated profiles retain their existing context, slot, batch, and KV
101
+ settings. Reasoning and speculation are independent, explicit controls layered
102
+ on those resource profiles:
103
+
104
+ | Setting | Behavior |
105
+ | --- | --- |
106
+ | `TREEBEARD_REASONING=off` | Default. Disables thinking in the server template while preserving explicit per-request overrides. |
107
+ | `TREEBEARD_REASONING=bounded` | Enables thinking with a default 64-token GPU or 16-token CPU budget. |
108
+ | `TREEBEARD_REASONING=unrestricted` | Enables thinking without a token budget. |
109
+ | `TREEBEARD_SPECULATION=off` | Default. Leaves the runtime's no-speculation default unchanged. |
110
+ | `TREEBEARD_SPECULATION=ngram` | Conservative `ngram-map-k` prompt-reuse drafting. |
111
+ | `TREEBEARD_SPECULATION=mtp` | Conservative two-token drafting with the model's native MTP head. |
112
+ | `TREEBEARD_SPECULATION=hybrid` | Tries n-gram drafting first, then native MTP. |
113
+
114
+ Set `TREEBEARD_REASONING_BUDGET` to a positive integer to override the bounded
115
+ default. The off setting is the server default; API clients can still opt an
116
+ individual request into thinking with request-level chat-template and budget
117
+ controls. Qwen3.6's native one-layer MTP head is carried by the GGUF, so `mtp`
118
+ and `hybrid` do not require another model. Additional `llama-server` arguments
119
+ may still be appended after `treebeard serve` for controlled experiments.
120
+
121
+ On the pinned b9624 runtime, selective OpenAI-compatible thinking must set both
122
+ `chat_template_kwargs.enable_thinking=true` and `thinking_budget_tokens=N` on
123
+ the request while the launcher remains in its default `off` mode. The request's
124
+ `max_tokens` limit includes both thought and answer tokens, and every model
125
+ turn after a tool result starts with a fresh budget. Global bounded reasoning
126
+ works on b9624, but overriding that global budget with a smaller request budget
127
+ requires the newer request-precedence fix. The packaged RC3 binaries also
128
+ predate newer Anthropic thinking-control translations; they require a runtime
129
+ rebuild and are not claimed by this launcher-only change.
130
+
131
+ The package makes no default speculative speed claim. N-gram hit rate, MTP
132
+ acceptance, verification cost, memory pressure, and reasoning quality are
133
+ workload-dependent and need matched evaluation before deployment.
run.sh CHANGED
@@ -2,7 +2,7 @@
2
  set -Eeuo pipefail
3
 
4
  VERSION=0.1.0-rc.3
5
- PACKAGING_REVISION=pkg3
6
  MODEL_NAME=Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
7
  MODEL_SHA256=25233af7642e3a91bd52cc4aeefdbd4a117479088e06cf1aea5b6bedb443c506
8
  MODEL_SIZE=26592508896
@@ -257,6 +257,15 @@ HOST="${TREEBEARD_HOST:-127.0.0.1}"
257
  PORT="${TREEBEARD_PORT:-8093}"
258
  MULTIMODAL="${TREEBEARD_MULTIMODAL:-0}"
259
  DRY_RUN="${TREEBEARD_DRY_RUN:-0}"
 
 
 
 
 
 
 
 
 
260
 
261
  require_positive_integer TREEBEARD_CONTEXT "$CONTEXT"
262
  require_positive_integer TREEBEARD_PARALLEL "$PARALLEL"
@@ -272,6 +281,25 @@ require_positive_integer TREEBEARD_PORT "$PORT"
272
  [[ "$FLASH_ATTN" == on || "$FLASH_ATTN" == off || "$FLASH_ATTN" == auto ]] ||
273
  fail "TREEBEARD_FLASH_ATTN must be on, off, or auto"
274
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
275
  mkdir -p "$CACHE_ROOT"
276
  chmod 700 "$CACHE_ROOT"
277
  verify_payload "$MODEL" "$MODEL_SHA256" "$MODEL_SIZE" "model"
@@ -325,6 +353,63 @@ args+=(
325
  -a "treebeard-$VERSION-Qwen3.6-35B-A3B-Q5-${BACKEND}-c${CONTEXT}-np${PARALLEL}"
326
  )
327
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
328
  if [[ "$BACKEND" == sycl ]]; then
329
  args+=(--no-op-offload)
330
  fi
@@ -333,8 +418,9 @@ if [[ "$MULTIMODAL" == 1 ]]; then
333
  fi
334
  args+=("$@")
335
 
336
- printf 'Treebeard %s %s: backend=%s profile=%s context=%s slots=%s multimodal=%s\n' \
337
- "$VERSION" "$PACKAGING_REVISION" "$BACKEND" "$PROFILE" "$CONTEXT" "$PARALLEL" "$MULTIMODAL"
 
338
 
339
  if [[ "$DRY_RUN" == 1 ]]; then
340
  printf 'Command:'
 
2
  set -Eeuo pipefail
3
 
4
  VERSION=0.1.0-rc.3
5
+ PACKAGING_REVISION=pkg4
6
  MODEL_NAME=Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
7
  MODEL_SHA256=25233af7642e3a91bd52cc4aeefdbd4a117479088e06cf1aea5b6bedb443c506
8
  MODEL_SIZE=26592508896
 
257
  PORT="${TREEBEARD_PORT:-8093}"
258
  MULTIMODAL="${TREEBEARD_MULTIMODAL:-0}"
259
  DRY_RUN="${TREEBEARD_DRY_RUN:-0}"
260
+ REASONING="${TREEBEARD_REASONING:-off}"
261
+ SPECULATION="${TREEBEARD_SPECULATION:-off}"
262
+
263
+ if [[ "$BACKEND" == cpu ]]; then
264
+ DEFAULT_REASONING_BUDGET=16
265
+ else
266
+ DEFAULT_REASONING_BUDGET=64
267
+ fi
268
+ REASONING_BUDGET="${TREEBEARD_REASONING_BUDGET:-$DEFAULT_REASONING_BUDGET}"
269
 
270
  require_positive_integer TREEBEARD_CONTEXT "$CONTEXT"
271
  require_positive_integer TREEBEARD_PARALLEL "$PARALLEL"
 
281
  [[ "$FLASH_ATTN" == on || "$FLASH_ATTN" == off || "$FLASH_ATTN" == auto ]] ||
282
  fail "TREEBEARD_FLASH_ATTN must be on, off, or auto"
283
 
284
+ case "$REASONING" in
285
+ off|unrestricted)
286
+ ;;
287
+ bounded)
288
+ require_positive_integer TREEBEARD_REASONING_BUDGET "$REASONING_BUDGET"
289
+ ;;
290
+ *)
291
+ fail "TREEBEARD_REASONING must be off, bounded, or unrestricted"
292
+ ;;
293
+ esac
294
+
295
+ case "$SPECULATION" in
296
+ off|ngram|mtp|hybrid)
297
+ ;;
298
+ *)
299
+ fail "TREEBEARD_SPECULATION must be off, ngram, mtp, or hybrid"
300
+ ;;
301
+ esac
302
+
303
  mkdir -p "$CACHE_ROOT"
304
  chmod 700 "$CACHE_ROOT"
305
  verify_payload "$MODEL" "$MODEL_SHA256" "$MODEL_SIZE" "model"
 
353
  -a "treebeard-$VERSION-Qwen3.6-35B-A3B-Q5-${BACKEND}-c${CONTEXT}-np${PARALLEL}"
354
  )
355
 
356
+ case "$REASONING" in
357
+ off)
358
+ # Preserve the validated no-thinking server default. Clients can still
359
+ # opt individual requests into thinking with request-level controls.
360
+ args+=(--reasoning off --reasoning-budget -1)
361
+ REASONING_DETAIL=off
362
+ ;;
363
+ bounded)
364
+ args+=(--reasoning on --reasoning-budget "$REASONING_BUDGET")
365
+ REASONING_DETAIL="$REASONING_BUDGET"
366
+ ;;
367
+ unrestricted)
368
+ args+=(--reasoning on --reasoning-budget -1)
369
+ REASONING_DETAIL=unlimited
370
+ ;;
371
+ esac
372
+
373
+ case "$SPECULATION" in
374
+ off)
375
+ # The pinned runtime already defaults to one NONE sentinel. Do not add
376
+ # another --spec-type none entry on its append-style parser.
377
+ ;;
378
+ ngram)
379
+ # Short drafts and two required prompt hits are deliberately
380
+ # conservative; acceptance and speed remain workload-dependent.
381
+ args+=(
382
+ --spec-type ngram-map-k
383
+ --spec-ngram-map-k-size-n 12
384
+ --spec-ngram-map-k-size-m 16
385
+ --spec-ngram-map-k-min-hits 2
386
+ )
387
+ ;;
388
+ mtp|hybrid)
389
+ spec_types=draft-mtp
390
+ if [[ "$SPECULATION" == hybrid ]]; then
391
+ spec_types=ngram-map-k,draft-mtp
392
+ args+=(
393
+ --spec-ngram-map-k-size-n 12
394
+ --spec-ngram-map-k-size-m 16
395
+ --spec-ngram-map-k-min-hits 2
396
+ )
397
+ fi
398
+ # Qwen3.6's native one-layer MTP head is carried by the model GGUF.
399
+ # Keep the verification batch narrow until a workload proves a larger
400
+ # draft profitable on its backend and concurrency shape.
401
+ args+=(
402
+ --spec-type "$spec_types"
403
+ --spec-draft-n-max 2
404
+ --spec-draft-n-min 1
405
+ --spec-draft-p-min 0.20
406
+ --spec-draft-mtp-branch-k 1
407
+ --spec-draft-mtp-tree-width 1
408
+ --spec-draft-mtp-tree-depth 2
409
+ )
410
+ ;;
411
+ esac
412
+
413
  if [[ "$BACKEND" == sycl ]]; then
414
  args+=(--no-op-offload)
415
  fi
 
418
  fi
419
  args+=("$@")
420
 
421
+ printf 'Treebeard %s %s: backend=%s profile=%s context=%s slots=%s multimodal=%s reasoning=%s(%s) speculation=%s\n' \
422
+ "$VERSION" "$PACKAGING_REVISION" "$BACKEND" "$PROFILE" "$CONTEXT" "$PARALLEL" "$MULTIMODAL" \
423
+ "$REASONING" "$REASONING_DETAIL" "$SPECULATION"
424
 
425
  if [[ "$DRY_RUN" == 1 ]]; then
426
  printf 'Command:'
treebeard CHANGED
@@ -22,6 +22,9 @@ Usage:
22
  Common settings:
23
  TREEBEARD_BACKEND=auto|cpu|sycl|cuda
24
  TREEBEARD_PROFILE=quality|throughput|custom
 
 
 
25
  TREEBEARD_CONTEXT=<tokens>
26
  TREEBEARD_HOST=127.0.0.1
27
  TREEBEARD_PORT=8093
 
22
  Common settings:
23
  TREEBEARD_BACKEND=auto|cpu|sycl|cuda
24
  TREEBEARD_PROFILE=quality|throughput|custom
25
+ TREEBEARD_REASONING=off|bounded|unrestricted
26
+ TREEBEARD_REASONING_BUDGET=<tokens> Bounded default: GPU 64, CPU 16
27
+ TREEBEARD_SPECULATION=off|ngram|mtp|hybrid
28
  TREEBEARD_CONTEXT=<tokens>
29
  TREEBEARD_HOST=127.0.0.1
30
  TREEBEARD_PORT=8093