Instructions to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL # Run inference directly in the terminal: llama cli -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL # Run inference directly in the terminal: llama cli -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL # Run inference directly in the terminal: ./llama-cli -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Use Docker
docker model run hf.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
- LM Studio
- Jan
- vLLM
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
- Ollama
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Ollama:
ollama run hf.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
- Unsloth Desktop
- Pi
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Docker Model Runner:
docker model run hf.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
- Lemonade
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Run and chat with the model
lemonade run user.Treebeard-Qwen3.6-35B-A3B-GGUF-UD-Q5_K_XL
List all available models
lemonade list
- Hermes Agent
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download docs/BENCHMARKS.md from Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 2.98 kB
-
https://huggingface.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF/resolve/main/docs/BENCHMARKS.md
- Command line
-
hf download hf://Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF/docs/BENCHMARKS.md
-
curl -L -o BENCHMARKS.md https://huggingface.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF/resolve/main/docs/BENCHMARKS.md
Benchmarks
Model package: https://huggingface.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF
GitHub: https://github.com/newjordan/treebeard
Site: https://newjordan.github.io/treebeard/
Public result index: https://github.com/newjordan/treebeard/tree/main/results
Agent Bench freeze (named run)
Score 94/100 (130/138 points):
- 69/69 scenarios completed;
- 63 pass, 4 partial, 2 fail;
- zero request errors;
- one server slot and one benchmark worker;
- 262,144 total context tokens;
- temperature 0, thinking disabled, seed 42;
- tool-eval-bench 2.1.0 at
8b3259b; - llama.cpp build
b9624-0424f677f(freeze pin).
Same score and outcome vector on Intel Arc Pro B70 and NVIDIA GB10.
Packaged evidence:
evidence/agent/single-slot-94/result.jsonevidence/agent/single-slot-94/index.htmlevidence/agent/single-slot-94/guard.logevidence/agent/single-slot-94/nvidia-result.jsonevidence/nvidia/agent-single-slot-94/result.json
Control A/B (2026-07-28, stock Q5, Arc Pro B70)
Same stock Q5 weights, np=1, c=262144, temp 0, seed 42, tool-eval-bench 2.1.0
@ 8b3259b. Clean upstream SYCL versus Treebeard package binary + package env.
| Arm | Public 69 | Held-out ho-pack-v1.1 |
|---|---|---|
| Original Qwen control | 91/100 (125/138) | 42/46 (91.3%) |
| Treebeard package | 91/100 (126/138) | 42/46 (91.3%) |
Sequential tg_p50 (np=1, c=32768, 5 prompts × 2): control 77.1 → Treebeard 88.9. Multi-slot p50/agent (np=12, n_predict=96): control 6.88 → Treebeard 26.33 (concurrent capacity).
Evidence (GitHub, not duplicated as GGUF LFS here):
- https://github.com/newjordan/treebeard/tree/main/results/private-verification-20260728/agent-bench-ab-20260728T220230Z
- https://github.com/newjordan/treebeard/blob/main/docs/RELEASE-20260728.md
Portable CPU package smoke
Independent install on AMD Ryzen 9 5950X:
- build identity:
b9624-6a6dc2def-cpu; - context: 4,096 tokens for the bounded smoke;
- chat output: exact
TREEBEARD READY; - chat generation: 9.302 tok/s;
- tool call: exact
multiply({"a":17,"b":23}); - tool-call generation: 7.400 tok/s;
- loader warnings, request errors, and assertion failures: zero.
Functional fallback only. Raw files under evidence/cpu-linux-x86_64.
SYCL serving
Released 12-slot profile: 194.023 aggregate tok/s. 8-slot: 182.005 aggregate
tok/s. Single-session aggregate flat at -0.121% vs the prior RC2 reference in
the package notes. Supporting evidence under evidence/sycl.
NVIDIA
GB10, compute capability 12.1, CUDA 13.3, Linux ARM64:
- CUDA correctness: 1,104/1,104
MUL_MATand 796/796MUL_MAT_ID; - Q8_0 direct 12-column latency: 5.22 to 3.97 us median (31.49% lower);
- Q8_0 MoE down latency: 463.06 to 445.20 us median (4.01% lower);
- native pp4096: 2,422.325 tok/s over five samples;
- native tg128: 59.614 tok/s over five samples;
- single-slot agent freeze: 94/100, 130/138, zero request errors.
See docs/NVIDIA.md and evidence/nvidia.