Image-Text-to-Text
GGUF
qwen3_5_moe
qwen3.6
llama.cpp
nvidia
cuda
sycl
intel-gpu
cpu
multimodal
tool-calling
long-context
conversational
imatrix
Instructions to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL # Run inference directly in the terminal: llama cli -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL # Run inference directly in the terminal: llama cli -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL # Run inference directly in the terminal: ./llama-cli -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Use Docker
docker model run hf.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
- LM Studio
- Jan
- vLLM
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
- Ollama
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Ollama:
ollama run hf.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
- Unsloth Desktop
- Pi
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Docker Model Runner:
docker model run hf.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
- Lemonade
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Run and chat with the model
lemonade run user.Treebeard-Qwen3.6-35B-A3B-GGUF-UD-Q5_K_XL
List all available models
lemonade list
- Hermes Agent
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Treebeard 0.1.1 launcher refresh (pkg4): reasoning + speculation controls
Browse filesrun.sh/treebeard pkg4: TREEBEARD_REASONING=off|bounded|unrestricted and
TREEBEARD_SPECULATION=off|ngram|mtp|hybrid; PACKAGE.json optional_controls;
card + PACKAGE.md docs; SHA256SUMS re-pinned. Runtime bytes unchanged.
- PACKAGE.json +26 -1
- README.md +14 -1
- SHA256SUMS +5 -5
- docs/PACKAGE.md +43 -0
- run.sh +89 -3
- treebeard +3 -0
PACKAGE.json
CHANGED
|
@@ -3,7 +3,7 @@
|
|
| 3 |
"name": "Treebeard-Qwen3.6-35B-A3B-GGUF",
|
| 4 |
"product": "Treebeard",
|
| 5 |
"version": "0.1.0-rc.3",
|
| 6 |
-
"packaging_revision": "
|
| 7 |
"visibility": "public",
|
| 8 |
"created_date": "2026-07-11",
|
| 9 |
"publication": {
|
|
@@ -103,6 +103,31 @@
|
|
| 103 |
"server_slots": 1
|
| 104 |
}
|
| 105 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 106 |
"agent_benchmark": {
|
| 107 |
"tool": "tool-eval-bench 2.1.0",
|
| 108 |
"commit": "8b3259be7411fe27c7610d0de64ae1d3b622b9ef",
|
|
|
|
| 3 |
"name": "Treebeard-Qwen3.6-35B-A3B-GGUF",
|
| 4 |
"product": "Treebeard",
|
| 5 |
"version": "0.1.0-rc.3",
|
| 6 |
+
"packaging_revision": "pkg4",
|
| 7 |
"visibility": "public",
|
| 8 |
"created_date": "2026-07-11",
|
| 9 |
"publication": {
|
|
|
|
| 103 |
"server_slots": 1
|
| 104 |
}
|
| 105 |
},
|
| 106 |
+
"optional_controls": {
|
| 107 |
+
"reasoning": {
|
| 108 |
+
"default": "off",
|
| 109 |
+
"modes": [
|
| 110 |
+
"off",
|
| 111 |
+
"bounded",
|
| 112 |
+
"unrestricted"
|
| 113 |
+
],
|
| 114 |
+
"bounded_budget_tokens": {
|
| 115 |
+
"gpu": 64,
|
| 116 |
+
"cpu": 16
|
| 117 |
+
}
|
| 118 |
+
},
|
| 119 |
+
"speculation": {
|
| 120 |
+
"default": "off",
|
| 121 |
+
"modes": [
|
| 122 |
+
"off",
|
| 123 |
+
"ngram",
|
| 124 |
+
"mtp",
|
| 125 |
+
"hybrid"
|
| 126 |
+
],
|
| 127 |
+
"native_mtp_layers": 1,
|
| 128 |
+
"validated_speedup": false
|
| 129 |
+
}
|
| 130 |
+
},
|
| 131 |
"agent_benchmark": {
|
| 132 |
"tool": "tool-eval-bench 2.1.0",
|
| 133 |
"commit": "8b3259be7411fe27c7610d0de64ae1d3b622b9ef",
|
README.md
CHANGED
|
@@ -21,7 +21,7 @@ pipeline_tag: image-text-to-text
|
|
| 21 |
# Treebeard
|
| 22 |
|
| 23 |
Treebeard is a ready-to-run Linux package for Qwen3.6-35B-A3B. Release
|
| 24 |
-
`0.1.0-rc.3` combines one 26.6 GB Q5_K_XL GGUF, the optional F16 vision
|
| 25 |
projector, official Qwen metadata and tokenizer files, platform runtimes, a
|
| 26 |
one-command installer, an OpenAI-compatible server launcher, and the raw
|
| 27 |
evidence behind its benchmark claims.
|
|
@@ -152,6 +152,9 @@ tokens. Override settings with environment variables:
|
|
| 152 |
| `TREEBEARD_HOST` | `127.0.0.1` | API bind address |
|
| 153 |
| `TREEBEARD_PORT` | `8093` | API port |
|
| 154 |
| `TREEBEARD_MULTIMODAL` | `0` | Load the installed F16 projector |
|
|
|
|
|
|
|
|
|
|
| 155 |
| `TREEBEARD_VERIFY` | `once` | `once`, `always`, or `never` |
|
| 156 |
| `TREEBEARD_CPUSET` | unset | Optional `taskset` CPU list |
|
| 157 |
|
|
@@ -168,6 +171,16 @@ not the configuration behind the single-slot 94:
|
|
| 168 |
TREEBEARD_PROFILE=throughput treebeard serve
|
| 169 |
```
|
| 170 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 171 |
## Full repository package
|
| 172 |
|
| 173 |
The Hugging Face repository is the complete package. After downloading it with
|
|
|
|
| 21 |
# Treebeard
|
| 22 |
|
| 23 |
Treebeard is a ready-to-run Linux package for Qwen3.6-35B-A3B. Release
|
| 24 |
+
`0.1.1` (runtime `0.1.0-rc.3`, launcher pkg4) combines one 26.6 GB Q5_K_XL GGUF, the optional F16 vision
|
| 25 |
projector, official Qwen metadata and tokenizer files, platform runtimes, a
|
| 26 |
one-command installer, an OpenAI-compatible server launcher, and the raw
|
| 27 |
evidence behind its benchmark claims.
|
|
|
|
| 152 |
| `TREEBEARD_HOST` | `127.0.0.1` | API bind address |
|
| 153 |
| `TREEBEARD_PORT` | `8093` | API port |
|
| 154 |
| `TREEBEARD_MULTIMODAL` | `0` | Load the installed F16 projector |
|
| 155 |
+
| `TREEBEARD_REASONING` | `off` | `off`, `bounded`, or `unrestricted` |
|
| 156 |
+
| `TREEBEARD_REASONING_BUDGET` | `64` GPU / `16` CPU | Bounded thinking tokens |
|
| 157 |
+
| `TREEBEARD_SPECULATION` | `off` | `off`, `ngram`, `mtp`, or `hybrid` |
|
| 158 |
| `TREEBEARD_VERIFY` | `once` | `once`, `always`, or `never` |
|
| 159 |
| `TREEBEARD_CPUSET` | unset | Optional `taskset` CPU list |
|
| 160 |
|
|
|
|
| 171 |
TREEBEARD_PROFILE=throughput treebeard serve
|
| 172 |
```
|
| 173 |
|
| 174 |
+
Reasoning is off by default, matching the validated 94/100 agent benchmark.
|
| 175 |
+
`TREEBEARD_REASONING=bounded` enables a small thinking allowance (64 tokens
|
| 176 |
+
on GPU, 16 on CPU; override with `TREEBEARD_REASONING_BUDGET`), and
|
| 177 |
+
`unrestricted` removes the budget. API clients can still opt individual
|
| 178 |
+
requests into thinking with request-level chat-template and thinking-budget
|
| 179 |
+
controls. Speculative decoding is opt-in through
|
| 180 |
+
`TREEBEARD_SPECULATION=ngram|mtp|hybrid` (`mtp` uses Qwen3.6's native
|
| 181 |
+
one-layer MTP head carried by the GGUF); these modes ship in the 0.1.1
|
| 182 |
+
launcher refresh but are not yet validated as a Treebeard speedup.
|
| 183 |
+
|
| 184 |
## Full repository package
|
| 185 |
|
| 186 |
The Hugging Face repository is the complete package. After downloading it with
|
SHA256SUMS
CHANGED
|
@@ -2,15 +2,15 @@ cae3aa82fc25bb3f9f6ae4c826f8c08dea0bba30658c448a7f124737549ce345 .gitattributes
|
|
| 2 |
50cbab8a892c5f2993b8c7351a99182507472def3b1374558308605d99b86b32 LICENSE
|
| 3 |
94f29bbed6a22c35b992c5c6ebf0e7c92f13b836b90f36f461c9cf2f0f1d010d LICENSE-RUNTIME
|
| 4 |
2d04af27c51a8ba55b9f0cbfcf1df72270daf70cf5e5f6d6cbd2db794e9c2151 NOTICE.md
|
| 5 |
-
|
| 6 |
25233af7642e3a91bd52cc4aeefdbd4a117479088e06cf1aea5b6bedb443c506 Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
|
| 7 |
-
|
| 8 |
e84f32a23fdda27689f868aa4a1a5621f41133e51a48d7f3efcbea2839574259 chat_template.jinja
|
| 9 |
93a4693fa9d8392fbfccd4b3c9873f4bfdcb14fdede978b123d07d19675efe99 config.json
|
| 10 |
c1b09db419119513247e9b8b912c4b9897106c9b20c6cada7e107d993c5435eb configuration.json
|
| 11 |
ae1c5fb00a620a316159929e229e4ca6a170efa6317e55164e614b120af417fd docs/BENCHMARKS.md
|
| 12 |
f648d15c3c111b72f775f9f08a43611d8398a9ff1855aba86d50576096a8b0c8 docs/NVIDIA.md
|
| 13 |
-
|
| 14 |
0daf18022e86b6e12510735aefa8ec52fea1c48b2f5c1136608287d48bda3f99 evidence/README.md
|
| 15 |
efc91f319aef7a32c311a4d0f058dd87c92c1b96352fc322966c8ce2ed6fd4eb evidence/agent/single-slot-94/LEGACY-RUN-IDENTITY.md
|
| 16 |
6e3b2f8e317b123ae8fc35d91ca44a3765fdba1a8d50502be59624a6e24d3809 evidence/agent/single-slot-94/SHA256SUMS
|
|
@@ -143,14 +143,14 @@ a9d356d7bdf1ef4949e3e748e95b8e10ad9d4e2e838eddc38a0a7b6b94d1db8d merges.txt
|
|
| 143 |
27225450ac9c6529872ee1924fcb0962ff5634834f817040f444118116f4e516 preprocessor_config.json
|
| 144 |
bcbcf46eae0ada57cfa71645004fd0850781ef2bad7c9da57f7b922169baa421 profiles/quality.env
|
| 145 |
97afe010ba8d3292456f875d115f03a257a95c12ae02d145c2bcaef449f4718e profiles/throughput.env
|
| 146 |
-
|
| 147 |
48f8c4abb9467fcd0e4c0e8284251903f0e40b27dbf20786f3d9969478c0c81a runtime/SHA256SUMS
|
| 148 |
51497d7a72cfe8cdf3d66fcbf3ba36f8d5130084cc339cdc49995ec131126732 runtime/cpu-linux-x86_64/treebeard-0.1.0-rc.3-b9624-pkg3-cpu-ubuntu22.04-linux-x86_64.tar.gz
|
| 149 |
0dfadb9ed53c66ab5a23b6731988d833b223af4e51c8f11ad9522a95eb0f04f3 runtime/cuda-linux-aarch64/treebeard-0.1.0-rc.3-b9624-pkg2-cuda13.3-linux-aarch64.tar.gz
|
| 150 |
beec6bd316cc5285481d1545e66b31f984b1eb93e69fe05c7d0b1c5e1bafc756 runtime/sycl-linux-x86_64/treebeard-0.1.0-rc.3-b9624-pkg2-sycl-oneapi2026-linux-x86_64.tar.gz
|
| 151 |
5f9e4d4901a92b997e463c1f46055088b6cca5ca61a6522d1b9f64c4bb81cb42 tokenizer.json
|
| 152 |
5186f0defcd7f232382c7f0aebcd2252d073bb921ab240e407b7ae8745d2b29b tokenizer_config.json
|
| 153 |
-
|
| 154 |
3660ebd0bf0e9ce8c3a73af312077029793c3d9a8f52d973c7cd5fc9ca5e5f63 verify.sh
|
| 155 |
7768af27c1fafa9cc9011c1dc20067e03f8915e03b63504550e11d5066986d13 video_preprocessor_config.json
|
| 156 |
ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003 vocab.json
|
|
|
|
| 2 |
50cbab8a892c5f2993b8c7351a99182507472def3b1374558308605d99b86b32 LICENSE
|
| 3 |
94f29bbed6a22c35b992c5c6ebf0e7c92f13b836b90f36f461c9cf2f0f1d010d LICENSE-RUNTIME
|
| 4 |
2d04af27c51a8ba55b9f0cbfcf1df72270daf70cf5e5f6d6cbd2db794e9c2151 NOTICE.md
|
| 5 |
+
7bdb30865738053188af49629d2fd9c0fbe9843d7c913727df01c7e554214f69 PACKAGE.json
|
| 6 |
25233af7642e3a91bd52cc4aeefdbd4a117479088e06cf1aea5b6bedb443c506 Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
|
| 7 |
+
a755e93db87862c16b9765b97a3bb9e2617d767a2aaa8565699e24a0257fa9aa README.md
|
| 8 |
e84f32a23fdda27689f868aa4a1a5621f41133e51a48d7f3efcbea2839574259 chat_template.jinja
|
| 9 |
93a4693fa9d8392fbfccd4b3c9873f4bfdcb14fdede978b123d07d19675efe99 config.json
|
| 10 |
c1b09db419119513247e9b8b912c4b9897106c9b20c6cada7e107d993c5435eb configuration.json
|
| 11 |
ae1c5fb00a620a316159929e229e4ca6a170efa6317e55164e614b120af417fd docs/BENCHMARKS.md
|
| 12 |
f648d15c3c111b72f775f9f08a43611d8398a9ff1855aba86d50576096a8b0c8 docs/NVIDIA.md
|
| 13 |
+
999acc5f1949ed80f5775503821b7f81ab1c31b47581e0ab884139dc49435398 docs/PACKAGE.md
|
| 14 |
0daf18022e86b6e12510735aefa8ec52fea1c48b2f5c1136608287d48bda3f99 evidence/README.md
|
| 15 |
efc91f319aef7a32c311a4d0f058dd87c92c1b96352fc322966c8ce2ed6fd4eb evidence/agent/single-slot-94/LEGACY-RUN-IDENTITY.md
|
| 16 |
6e3b2f8e317b123ae8fc35d91ca44a3765fdba1a8d50502be59624a6e24d3809 evidence/agent/single-slot-94/SHA256SUMS
|
|
|
|
| 143 |
27225450ac9c6529872ee1924fcb0962ff5634834f817040f444118116f4e516 preprocessor_config.json
|
| 144 |
bcbcf46eae0ada57cfa71645004fd0850781ef2bad7c9da57f7b922169baa421 profiles/quality.env
|
| 145 |
97afe010ba8d3292456f875d115f03a257a95c12ae02d145c2bcaef449f4718e profiles/throughput.env
|
| 146 |
+
acd5c33ea206ba86814c5c739647a08f80b693332354c7a753a1eac9bade9206 run.sh
|
| 147 |
48f8c4abb9467fcd0e4c0e8284251903f0e40b27dbf20786f3d9969478c0c81a runtime/SHA256SUMS
|
| 148 |
51497d7a72cfe8cdf3d66fcbf3ba36f8d5130084cc339cdc49995ec131126732 runtime/cpu-linux-x86_64/treebeard-0.1.0-rc.3-b9624-pkg3-cpu-ubuntu22.04-linux-x86_64.tar.gz
|
| 149 |
0dfadb9ed53c66ab5a23b6731988d833b223af4e51c8f11ad9522a95eb0f04f3 runtime/cuda-linux-aarch64/treebeard-0.1.0-rc.3-b9624-pkg2-cuda13.3-linux-aarch64.tar.gz
|
| 150 |
beec6bd316cc5285481d1545e66b31f984b1eb93e69fe05c7d0b1c5e1bafc756 runtime/sycl-linux-x86_64/treebeard-0.1.0-rc.3-b9624-pkg2-sycl-oneapi2026-linux-x86_64.tar.gz
|
| 151 |
5f9e4d4901a92b997e463c1f46055088b6cca5ca61a6522d1b9f64c4bb81cb42 tokenizer.json
|
| 152 |
5186f0defcd7f232382c7f0aebcd2252d073bb921ab240e407b7ae8745d2b29b tokenizer_config.json
|
| 153 |
+
16eb404820fb872f14460a546e79f3cf99c98d2314b40167c105ed08f2f1c132 treebeard
|
| 154 |
3660ebd0bf0e9ce8c3a73af312077029793c3d9a8f52d973c7cd5fc9ca5e5f63 verify.sh
|
| 155 |
7768af27c1fafa9cc9011c1dc20067e03f8915e03b63504550e11d5066986d13 video_preprocessor_config.json
|
| 156 |
ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003 vocab.json
|
docs/PACKAGE.md
CHANGED
|
@@ -1,5 +1,11 @@
|
|
| 1 |
# Package contract
|
| 2 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
Treebeard pkg3 follows a standard Hugging Face GGUF layout. The complete GGUF
|
| 4 |
model and official Qwen configuration and tokenizer files live at repository
|
| 5 |
root. Runtimes, launch tools, and evidence are additive directories.
|
|
@@ -88,3 +94,40 @@ correctness where applicable, and a real server/API package smoke test.
|
|
| 88 |
The model supports up to 262,144 context tokens, but actual usable context is
|
| 89 |
bounded by available device or system memory. The launcher exposes overrides
|
| 90 |
instead of claiming every host can sustain the maximum profile.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
# Package contract
|
| 2 |
|
| 3 |
+
Model package: <https://huggingface.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF>
|
| 4 |
+
|
| 5 |
+
GitHub repository: <https://github.com/newjordan/treebeard>
|
| 6 |
+
|
| 7 |
+
MoE algorithm explainer: <https://newjordan.github.io/treebeard/moe-routing.html>
|
| 8 |
+
|
| 9 |
Treebeard pkg3 follows a standard Hugging Face GGUF layout. The complete GGUF
|
| 10 |
model and official Qwen configuration and tokenizer files live at repository
|
| 11 |
root. Runtimes, launch tools, and evidence are additive directories.
|
|
|
|
| 94 |
The model supports up to 262,144 context tokens, but actual usable context is
|
| 95 |
bounded by available device or system memory. The launcher exposes overrides
|
| 96 |
instead of claiming every host can sustain the maximum profile.
|
| 97 |
+
|
| 98 |
+
## Reasoning and speculative decoding
|
| 99 |
+
|
| 100 |
+
The validated profiles retain their existing context, slot, batch, and KV
|
| 101 |
+
settings. Reasoning and speculation are independent, explicit controls layered
|
| 102 |
+
on those resource profiles:
|
| 103 |
+
|
| 104 |
+
| Setting | Behavior |
|
| 105 |
+
| --- | --- |
|
| 106 |
+
| `TREEBEARD_REASONING=off` | Default. Disables thinking in the server template while preserving explicit per-request overrides. |
|
| 107 |
+
| `TREEBEARD_REASONING=bounded` | Enables thinking with a default 64-token GPU or 16-token CPU budget. |
|
| 108 |
+
| `TREEBEARD_REASONING=unrestricted` | Enables thinking without a token budget. |
|
| 109 |
+
| `TREEBEARD_SPECULATION=off` | Default. Leaves the runtime's no-speculation default unchanged. |
|
| 110 |
+
| `TREEBEARD_SPECULATION=ngram` | Conservative `ngram-map-k` prompt-reuse drafting. |
|
| 111 |
+
| `TREEBEARD_SPECULATION=mtp` | Conservative two-token drafting with the model's native MTP head. |
|
| 112 |
+
| `TREEBEARD_SPECULATION=hybrid` | Tries n-gram drafting first, then native MTP. |
|
| 113 |
+
|
| 114 |
+
Set `TREEBEARD_REASONING_BUDGET` to a positive integer to override the bounded
|
| 115 |
+
default. The off setting is the server default; API clients can still opt an
|
| 116 |
+
individual request into thinking with request-level chat-template and budget
|
| 117 |
+
controls. Qwen3.6's native one-layer MTP head is carried by the GGUF, so `mtp`
|
| 118 |
+
and `hybrid` do not require another model. Additional `llama-server` arguments
|
| 119 |
+
may still be appended after `treebeard serve` for controlled experiments.
|
| 120 |
+
|
| 121 |
+
On the pinned b9624 runtime, selective OpenAI-compatible thinking must set both
|
| 122 |
+
`chat_template_kwargs.enable_thinking=true` and `thinking_budget_tokens=N` on
|
| 123 |
+
the request while the launcher remains in its default `off` mode. The request's
|
| 124 |
+
`max_tokens` limit includes both thought and answer tokens, and every model
|
| 125 |
+
turn after a tool result starts with a fresh budget. Global bounded reasoning
|
| 126 |
+
works on b9624, but overriding that global budget with a smaller request budget
|
| 127 |
+
requires the newer request-precedence fix. The packaged RC3 binaries also
|
| 128 |
+
predate newer Anthropic thinking-control translations; they require a runtime
|
| 129 |
+
rebuild and are not claimed by this launcher-only change.
|
| 130 |
+
|
| 131 |
+
The package makes no default speculative speed claim. N-gram hit rate, MTP
|
| 132 |
+
acceptance, verification cost, memory pressure, and reasoning quality are
|
| 133 |
+
workload-dependent and need matched evaluation before deployment.
|
run.sh
CHANGED
|
@@ -2,7 +2,7 @@
|
|
| 2 |
set -Eeuo pipefail
|
| 3 |
|
| 4 |
VERSION=0.1.0-rc.3
|
| 5 |
-
PACKAGING_REVISION=
|
| 6 |
MODEL_NAME=Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
|
| 7 |
MODEL_SHA256=25233af7642e3a91bd52cc4aeefdbd4a117479088e06cf1aea5b6bedb443c506
|
| 8 |
MODEL_SIZE=26592508896
|
|
@@ -257,6 +257,15 @@ HOST="${TREEBEARD_HOST:-127.0.0.1}"
|
|
| 257 |
PORT="${TREEBEARD_PORT:-8093}"
|
| 258 |
MULTIMODAL="${TREEBEARD_MULTIMODAL:-0}"
|
| 259 |
DRY_RUN="${TREEBEARD_DRY_RUN:-0}"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 260 |
|
| 261 |
require_positive_integer TREEBEARD_CONTEXT "$CONTEXT"
|
| 262 |
require_positive_integer TREEBEARD_PARALLEL "$PARALLEL"
|
|
@@ -272,6 +281,25 @@ require_positive_integer TREEBEARD_PORT "$PORT"
|
|
| 272 |
[[ "$FLASH_ATTN" == on || "$FLASH_ATTN" == off || "$FLASH_ATTN" == auto ]] ||
|
| 273 |
fail "TREEBEARD_FLASH_ATTN must be on, off, or auto"
|
| 274 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 275 |
mkdir -p "$CACHE_ROOT"
|
| 276 |
chmod 700 "$CACHE_ROOT"
|
| 277 |
verify_payload "$MODEL" "$MODEL_SHA256" "$MODEL_SIZE" "model"
|
|
@@ -325,6 +353,63 @@ args+=(
|
|
| 325 |
-a "treebeard-$VERSION-Qwen3.6-35B-A3B-Q5-${BACKEND}-c${CONTEXT}-np${PARALLEL}"
|
| 326 |
)
|
| 327 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 328 |
if [[ "$BACKEND" == sycl ]]; then
|
| 329 |
args+=(--no-op-offload)
|
| 330 |
fi
|
|
@@ -333,8 +418,9 @@ if [[ "$MULTIMODAL" == 1 ]]; then
|
|
| 333 |
fi
|
| 334 |
args+=("$@")
|
| 335 |
|
| 336 |
-
printf 'Treebeard %s %s: backend=%s profile=%s context=%s slots=%s multimodal=%s\n' \
|
| 337 |
-
"$VERSION" "$PACKAGING_REVISION" "$BACKEND" "$PROFILE" "$CONTEXT" "$PARALLEL" "$MULTIMODAL"
|
|
|
|
| 338 |
|
| 339 |
if [[ "$DRY_RUN" == 1 ]]; then
|
| 340 |
printf 'Command:'
|
|
|
|
| 2 |
set -Eeuo pipefail
|
| 3 |
|
| 4 |
VERSION=0.1.0-rc.3
|
| 5 |
+
PACKAGING_REVISION=pkg4
|
| 6 |
MODEL_NAME=Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
|
| 7 |
MODEL_SHA256=25233af7642e3a91bd52cc4aeefdbd4a117479088e06cf1aea5b6bedb443c506
|
| 8 |
MODEL_SIZE=26592508896
|
|
|
|
| 257 |
PORT="${TREEBEARD_PORT:-8093}"
|
| 258 |
MULTIMODAL="${TREEBEARD_MULTIMODAL:-0}"
|
| 259 |
DRY_RUN="${TREEBEARD_DRY_RUN:-0}"
|
| 260 |
+
REASONING="${TREEBEARD_REASONING:-off}"
|
| 261 |
+
SPECULATION="${TREEBEARD_SPECULATION:-off}"
|
| 262 |
+
|
| 263 |
+
if [[ "$BACKEND" == cpu ]]; then
|
| 264 |
+
DEFAULT_REASONING_BUDGET=16
|
| 265 |
+
else
|
| 266 |
+
DEFAULT_REASONING_BUDGET=64
|
| 267 |
+
fi
|
| 268 |
+
REASONING_BUDGET="${TREEBEARD_REASONING_BUDGET:-$DEFAULT_REASONING_BUDGET}"
|
| 269 |
|
| 270 |
require_positive_integer TREEBEARD_CONTEXT "$CONTEXT"
|
| 271 |
require_positive_integer TREEBEARD_PARALLEL "$PARALLEL"
|
|
|
|
| 281 |
[[ "$FLASH_ATTN" == on || "$FLASH_ATTN" == off || "$FLASH_ATTN" == auto ]] ||
|
| 282 |
fail "TREEBEARD_FLASH_ATTN must be on, off, or auto"
|
| 283 |
|
| 284 |
+
case "$REASONING" in
|
| 285 |
+
off|unrestricted)
|
| 286 |
+
;;
|
| 287 |
+
bounded)
|
| 288 |
+
require_positive_integer TREEBEARD_REASONING_BUDGET "$REASONING_BUDGET"
|
| 289 |
+
;;
|
| 290 |
+
*)
|
| 291 |
+
fail "TREEBEARD_REASONING must be off, bounded, or unrestricted"
|
| 292 |
+
;;
|
| 293 |
+
esac
|
| 294 |
+
|
| 295 |
+
case "$SPECULATION" in
|
| 296 |
+
off|ngram|mtp|hybrid)
|
| 297 |
+
;;
|
| 298 |
+
*)
|
| 299 |
+
fail "TREEBEARD_SPECULATION must be off, ngram, mtp, or hybrid"
|
| 300 |
+
;;
|
| 301 |
+
esac
|
| 302 |
+
|
| 303 |
mkdir -p "$CACHE_ROOT"
|
| 304 |
chmod 700 "$CACHE_ROOT"
|
| 305 |
verify_payload "$MODEL" "$MODEL_SHA256" "$MODEL_SIZE" "model"
|
|
|
|
| 353 |
-a "treebeard-$VERSION-Qwen3.6-35B-A3B-Q5-${BACKEND}-c${CONTEXT}-np${PARALLEL}"
|
| 354 |
)
|
| 355 |
|
| 356 |
+
case "$REASONING" in
|
| 357 |
+
off)
|
| 358 |
+
# Preserve the validated no-thinking server default. Clients can still
|
| 359 |
+
# opt individual requests into thinking with request-level controls.
|
| 360 |
+
args+=(--reasoning off --reasoning-budget -1)
|
| 361 |
+
REASONING_DETAIL=off
|
| 362 |
+
;;
|
| 363 |
+
bounded)
|
| 364 |
+
args+=(--reasoning on --reasoning-budget "$REASONING_BUDGET")
|
| 365 |
+
REASONING_DETAIL="$REASONING_BUDGET"
|
| 366 |
+
;;
|
| 367 |
+
unrestricted)
|
| 368 |
+
args+=(--reasoning on --reasoning-budget -1)
|
| 369 |
+
REASONING_DETAIL=unlimited
|
| 370 |
+
;;
|
| 371 |
+
esac
|
| 372 |
+
|
| 373 |
+
case "$SPECULATION" in
|
| 374 |
+
off)
|
| 375 |
+
# The pinned runtime already defaults to one NONE sentinel. Do not add
|
| 376 |
+
# another --spec-type none entry on its append-style parser.
|
| 377 |
+
;;
|
| 378 |
+
ngram)
|
| 379 |
+
# Short drafts and two required prompt hits are deliberately
|
| 380 |
+
# conservative; acceptance and speed remain workload-dependent.
|
| 381 |
+
args+=(
|
| 382 |
+
--spec-type ngram-map-k
|
| 383 |
+
--spec-ngram-map-k-size-n 12
|
| 384 |
+
--spec-ngram-map-k-size-m 16
|
| 385 |
+
--spec-ngram-map-k-min-hits 2
|
| 386 |
+
)
|
| 387 |
+
;;
|
| 388 |
+
mtp|hybrid)
|
| 389 |
+
spec_types=draft-mtp
|
| 390 |
+
if [[ "$SPECULATION" == hybrid ]]; then
|
| 391 |
+
spec_types=ngram-map-k,draft-mtp
|
| 392 |
+
args+=(
|
| 393 |
+
--spec-ngram-map-k-size-n 12
|
| 394 |
+
--spec-ngram-map-k-size-m 16
|
| 395 |
+
--spec-ngram-map-k-min-hits 2
|
| 396 |
+
)
|
| 397 |
+
fi
|
| 398 |
+
# Qwen3.6's native one-layer MTP head is carried by the model GGUF.
|
| 399 |
+
# Keep the verification batch narrow until a workload proves a larger
|
| 400 |
+
# draft profitable on its backend and concurrency shape.
|
| 401 |
+
args+=(
|
| 402 |
+
--spec-type "$spec_types"
|
| 403 |
+
--spec-draft-n-max 2
|
| 404 |
+
--spec-draft-n-min 1
|
| 405 |
+
--spec-draft-p-min 0.20
|
| 406 |
+
--spec-draft-mtp-branch-k 1
|
| 407 |
+
--spec-draft-mtp-tree-width 1
|
| 408 |
+
--spec-draft-mtp-tree-depth 2
|
| 409 |
+
)
|
| 410 |
+
;;
|
| 411 |
+
esac
|
| 412 |
+
|
| 413 |
if [[ "$BACKEND" == sycl ]]; then
|
| 414 |
args+=(--no-op-offload)
|
| 415 |
fi
|
|
|
|
| 418 |
fi
|
| 419 |
args+=("$@")
|
| 420 |
|
| 421 |
+
printf 'Treebeard %s %s: backend=%s profile=%s context=%s slots=%s multimodal=%s reasoning=%s(%s) speculation=%s\n' \
|
| 422 |
+
"$VERSION" "$PACKAGING_REVISION" "$BACKEND" "$PROFILE" "$CONTEXT" "$PARALLEL" "$MULTIMODAL" \
|
| 423 |
+
"$REASONING" "$REASONING_DETAIL" "$SPECULATION"
|
| 424 |
|
| 425 |
if [[ "$DRY_RUN" == 1 ]]; then
|
| 426 |
printf 'Command:'
|
treebeard
CHANGED
|
@@ -22,6 +22,9 @@ Usage:
|
|
| 22 |
Common settings:
|
| 23 |
TREEBEARD_BACKEND=auto|cpu|sycl|cuda
|
| 24 |
TREEBEARD_PROFILE=quality|throughput|custom
|
|
|
|
|
|
|
|
|
|
| 25 |
TREEBEARD_CONTEXT=<tokens>
|
| 26 |
TREEBEARD_HOST=127.0.0.1
|
| 27 |
TREEBEARD_PORT=8093
|
|
|
|
| 22 |
Common settings:
|
| 23 |
TREEBEARD_BACKEND=auto|cpu|sycl|cuda
|
| 24 |
TREEBEARD_PROFILE=quality|throughput|custom
|
| 25 |
+
TREEBEARD_REASONING=off|bounded|unrestricted
|
| 26 |
+
TREEBEARD_REASONING_BUDGET=<tokens> Bounded default: GPU 64, CPU 16
|
| 27 |
+
TREEBEARD_SPECULATION=off|ngram|mtp|hybrid
|
| 28 |
TREEBEARD_CONTEXT=<tokens>
|
| 29 |
TREEBEARD_HOST=127.0.0.1
|
| 30 |
TREEBEARD_PORT=8093
|