Instructions to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL # Run inference directly in the terminal: llama cli -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL # Run inference directly in the terminal: llama cli -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL # Run inference directly in the terminal: ./llama-cli -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Use Docker
docker model run hf.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
- LM Studio
- Jan
- vLLM
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
- Ollama
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Ollama:
ollama run hf.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
- Unsloth Desktop
- Pi
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Docker Model Runner:
docker model run hf.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
- Lemonade
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Run and chat with the model
lemonade run user.Treebeard-Qwen3.6-35B-A3B-GGUF-UD-Q5_K_XL
List all available models
lemonade list
- Hermes Agent
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF:UD-Q5_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download docs/PACKAGE.md from Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 6.54 kB
-
https://huggingface.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF/resolve/main/docs/PACKAGE.md
- Command line
-
hf download hf://Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF/docs/PACKAGE.md
-
curl -L -o PACKAGE.md https://huggingface.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF/resolve/main/docs/PACKAGE.md
Package contract
Model package: https://huggingface.co/Frosty40/Treebeard-Qwen3.6-35B-A3B-GGUF
GitHub repository: https://github.com/newjordan/treebeard
MoE algorithm explainer: https://newjordan.github.io/treebeard/moe-routing.html
Treebeard pkg3 follows a standard Hugging Face GGUF layout. The complete GGUF model and official Qwen configuration and tokenizer files live at repository root. Runtimes, launch tools, and evidence are additive directories.
User paths
There are two supported ways to use the package:
install.shdownloads the one model file and the single runtime selected for the host. This is the smallest and easiest installation path.- A full repository download includes every runtime, the optional projector, all standard model metadata, and all evidence. It is self-contained after download and does not fetch weights at launch time.
The installed treebeard command provides serve, doctor, verify,
status, report, and version commands. run.sh remains the direct launcher
and accepts additional llama-server arguments.
The CLI invokes packaged shell components through bash, so full-package
downloads remain usable when an archive client does not preserve executable
permission bits. bash ./treebeard serve is always a valid entry point.
Required root files
Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf;mmproj-F16.gguffor the full multimodal package;config.json,generation_config.json, andconfiguration.json;tokenizer.json,tokenizer_config.json,vocab.json, andmerges.txt;chat_template.jinja;preprocessor_config.jsonandvideo_preprocessor_config.json;install.sh,treebeard,run.sh, andverify.sh;PACKAGE.json,SHA256SUMS, andruntime/SHA256SUMS.
Runtime selection
run.sh and install.sh use the same ordered selection:
- CUDA on Linux ARM64 when a working NVIDIA device is present;
- SYCL on Linux x86_64 when a Level Zero GPU is present;
- the portable CPU runtime on Linux x86_64.
Explicit selection uses TREEBEARD_BACKEND=cpu|sycl|cuda. The CPU runtime was
built on Ubuntu 22.04 with GCC 11, glibc 2.35, GGML_NATIVE=OFF, dynamic backend
loading, and 14 x86_64 CPU variants. The runtime selects the best compatible
variant on the destination host.
The selected archive is SHA-256 verified and extracted under the user cache. Its internal manifest is checked on first extraction. The launcher never silently selects an unsupported architecture.
Verification
verify.sh checks the complete repository against SHA256SUMS. The installer
creates INSTALL-SHA256SUMS for the selected subset, and treebeard verify
checks the appropriate manifest automatically.
run.sh checks exact model size at each launch and verifies its hash once per
stable device/inode/size/mtime fingerprint by default. Set
TREEBEARD_VERIFY=always to rehash it on every launch. Runtime archives are
always hashed before extraction.
Compatibility
| Directory | Target | External requirement |
|---|---|---|
runtime/cpu-linux-x86_64 |
Linux x86_64 | glibc 2.35+ |
runtime/sycl-linux-x86_64 |
Linux x86_64 | Intel oneAPI 2026 and Level Zero |
runtime/cuda-linux-aarch64 |
Linux ARM64 | CUDA 13 runtime, cuBLAS, and compatible NVIDIA driver |
The CUDA and SYCL archives use vendor libraries installed on the host. GPU toolkits are not redistributed. NVIDIA x86_64 acceleration is not included in RC3; those systems select the portable CPU runtime.
Adding another target requires a runtime directory, an entry in
runtime/SHA256SUMS, build provenance, internal file hashes, CPU-oracle
correctness where applicable, and a real server/API package smoke test.
Resource contract
- text install download: about 26.7 GB;
- multimodal add-on: about 0.9 GB;
- minimum memory check: 32 GB, with an explicit low-memory override;
- default CPU profile: 32,768 context tokens and one slot;
- default validated GPU profile: 262,144 context tokens and one slot.
The model supports up to 262,144 context tokens, but actual usable context is bounded by available device or system memory. The launcher exposes overrides instead of claiming every host can sustain the maximum profile.
Reasoning and speculative decoding
The validated profiles retain their existing context, slot, batch, and KV settings. Reasoning and speculation are independent, explicit controls layered on those resource profiles:
| Setting | Behavior |
|---|---|
TREEBEARD_REASONING=off |
Default. Disables thinking in the server template while preserving explicit per-request overrides. |
TREEBEARD_REASONING=bounded |
Enables thinking with a default 64-token GPU or 16-token CPU budget. |
TREEBEARD_REASONING=unrestricted |
Enables thinking without a token budget. |
TREEBEARD_SPECULATION=off |
Default. Leaves the runtime's no-speculation default unchanged. |
TREEBEARD_SPECULATION=ngram |
Conservative ngram-map-k prompt-reuse drafting. |
TREEBEARD_SPECULATION=mtp |
Conservative two-token drafting with the model's native MTP head. |
TREEBEARD_SPECULATION=hybrid |
Tries n-gram drafting first, then native MTP. |
Set TREEBEARD_REASONING_BUDGET to a positive integer to override the bounded
default. The off setting is the server default; API clients can still opt an
individual request into thinking with request-level chat-template and budget
controls. Qwen3.6's native one-layer MTP head is carried by the GGUF, so mtp
and hybrid do not require another model. Additional llama-server arguments
may still be appended after treebeard serve for controlled experiments.
On the pinned b9624 runtime, selective OpenAI-compatible thinking must set both
chat_template_kwargs.enable_thinking=true and thinking_budget_tokens=N on
the request while the launcher remains in its default off mode. The request's
max_tokens limit includes both thought and answer tokens, and every model
turn after a tool result starts with a fresh budget. Global bounded reasoning
works on b9624, but overriding that global budget with a smaller request budget
requires the newer request-precedence fix. The packaged RC3 binaries also
predate newer Anthropic thinking-control translations; they require a runtime
rebuild and are not claimed by this launcher-only change.
The package makes no default speculative speed claim. N-gram hit rate, MTP acceptance, verification cost, memory pressure, and reasoning quality are workload-dependent and need matched evaluation before deployment.