Instructions to use prism-ml/Ternary-Bonsai-2-27B-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- LM Studio
- Jan
- vLLM
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prism-ml/Ternary-Bonsai-2-27B-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prism-ml/Ternary-Bonsai-2-27B-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Ollama
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Ollama:
ollama run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Unsloth Desktop
- Pi
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Docker Model Runner:
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Lemonade
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run and chat with the model
lemonade run user.Ternary-Bonsai-2-27B-gguf-F16
List all available models
lemonade list
- Hermes Agent
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
In-file MTP speculation refused on Hadamard-folded weights (prism-b10683); ~1.34x decode after the 7-line embedding fix
Hi β sharing a reproducible finding from running Bonsai 2 27B locally (RTX 4080 SUPER, Ada SM89, CUDA 13.3, WSL2). Short version: in-file MTP is a real ~1.34x decode win, but the official fork refuses the MTP graph on Hadamard-folded weights; a 7-line embedding-boundary fix (already in the community runtime) makes it work.
What works: MTP block packaged inside the main GGUF
ProCreations/Ternary-Bonsai-2-27B-MTP = the unchanged PQ2_0 target + blk.64.nextn.* (nextn_predict_layers=1) in one file. Launching with -m <bundle> --spec-type draft-mtp (no -md) builds the draft context against the target model (log: creating MTP draft context against the target model), so the draft shares the target's token_embd / output instead of carrying its own 248320x5120 copies.
That detail is the whole economics. A separate MTP sidecar carries a duplicate vocabulary β ~92% of its bytes β which puts the per-token draft cost ratio at rho = draft/target ~ 0.43 and makes even 40% acceptance a net loss (I measured 5.16 t/s vs 16.6 t/s no-spec on a 27B sidecar). Sharing the vocabulary drops rho to ~0.06 and the same head becomes profitable. Speedup ~ (1-a^(k+1))/(1-a) / (k*rho + 1).
What fails on the official fork
With prism-b10683-d8f26ee the same bundle refuses to start:
E llama_init_from_model: failed to initialize the context: Hadamard-latent table 'token_embd.weight' is read without the inverse transform
E common_speculative_init_result: failed to create MTP context
The guard is llama_verify_hadamard_graph; both messages are visible via strings libllama.so, and I could not find any flag to bypass it. It looks intentional: the MTP graph does ggml_get_rows(tok_embd_w, inp->tokens) on a table stored in the rotated basis.
The fix
Apply the inverse transform (rotation + explicit signs) to the MTP embedding lookup, i.e. 7 lines in src/models/qwen35.cpp::graph_mtp immediately after the ggml_get_rows. The community runtime ships exactly this as bonsai-mtp-embedding.patch, and its runtime/manifest.json says the bundled source (prism-dflash2-source.tar.gz, base d8f26ee) is already patched. I built it locally for SM89 (their prebuilt binary is SM120-only):
cmake -S llama -B llama/build -G Ninja -DGGML_CUDA=ON \
-DCMAKE_CUDA_COMPILER=/usr/local/cuda/bin/nvcc \
-DCMAKE_CUDA_ARCHITECTURES=89 -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF
(note: with nvcc not on PATH, CMake fails with CMAKE_CUDA_COMPILER-NOTFOUND unless the compiler is passed explicitly.)
Measured (vs --spec-type none, identical flags otherwise)
-ngl 99 -fa on -c 262144 -ctk q4_0 -ctv q4_0, temp 0, top_k 1, seed 7, n=128, 12 prompts, draft length 2:
| category (n) | baseline t/s | MTP t/s | speedup | acceptance |
|---|---|---|---|---|
| reasoning prose (3) | 66.4 | 84.9 | 1.25-1.31x | 47.7-57.6% |
| code continuation (3) | 67.4 | 101.3 | 1.32-1.70x | 62.5-94.3% |
| math, step-by-step (2) | 68.2 | 90.6 | 1.23-1.43x | 54.5-72.1% |
| format / repetitive (2) | 68.3 | 111.9 | 1.63-1.64x | 90.0-92.1% |
| Chinese (1) | 68.5 | 91.7 | 1.34x | 64.5% |
| median | 67.5 | 90.4 | 1.34x | 68.1% (803/1180) |
Same configuration, prefill of a ~12.5k-token prompt: 1726 t/s with PQ2_0 vs 1073 t/s with PTQ1_0 at 262144 + q4_0 KV; VRAM 15.7 / 16.0 GiB. (Consistent with the model card's "PQ2_0 wins prefill, PTQ1_0 wins decode on Ada".)
Suggestion
Please consider upstreaming the MTP embedding inverse-transform (and auditing any other latent table the draft graph reads) so that in-file MTP works on the official binaries. Today it only works with the community-patched runtime, which is SM120-only for the prebuilt archive and therefore forces non-Blackwell users to build from source. Happy to provide the patch, logs, or a repro script.
Adding the raw numbers behind the summary table above, plus a public reproduction bundle.
Public reproduction bundle: https://huggingface.co/datasets/zhaokeqi/bonsai2-27b-mtp-repro
It contains the raw per-prompt JSON for both arms (with the full timings objects), the A/B harness, the driver scripts and these notes, so the numbers above can be checked rather than trusted.
Measurement protocol (identical in both arms except --spec-type):
llama-server -m Ternary-Bonsai-2-27B-PQ2_0-MTP-Q8_0.gguf \
-ngl 99 -fa on -c 262144 -ctk q4_0 -ctv q4_0 -np 1 -t 16 --temp 0 \
--spec-type none # arm A
--spec-type draft-mtp --spec-draft-n-max 2 # arm B
POST /completion, temperature=0, top_k=1, seed=7, cache_prompt=false, n_predict=128. Acceptance is read from the response timings.draft_n / draft_n_accepted β note that llama-cli does not print acceptance at all, only the server does (or the log line slot print_timing: draft acceptance = ...).
Per-prompt results (single run per prompt β no repeats, so read the medians, not the last digit):
| # | prompt | baseline t/s | draft-mtp t/s | speedup | acceptance |
|---|---|---|---|---|---|
| R1 | reasoning prose | 67.2 | 88.0 | 1.31x | 57.6% |
| R2 | reasoning prose | 65.4 | 82.0 | 1.25x | 47.7% |
| R3 | reasoning prose | 66.6 | 84.6 | 1.27x | 57.6% |
| C1 | Python continuation | 67.1 | 99.8 | 1.49x | 77.0% |
| C2 | Python continuation | 67.4 | 114.8 | 1.70x | 94.3% |
| C3 | async Python | 67.7 | 89.2 | 1.32x | 62.5% |
| M1 | step-by-step math | 68.1 | 97.2 | 1.43x | 72.1% |
| M2 | probability recursion | 68.2 | 84.0 | 1.23x | 54.5% |
| F1 | JSON repetition | 68.3 | 111.5 | 1.63x | 90.0% |
| F2 | list continuation | 68.3 | 112.3 | 1.64x | 92.1% |
| Z1 | Chinese rewrite | 68.5 | 91.7 | 1.34x | 64.5% |
| median | 67.5 | 90.4 | 1.338x | 68.1% (803/1180) |
Two practical notes for 16 GB cards
- At
-c 262144on a 16 GB Ada card, the KV cache type decides everything:q8_0makes prefill collapse to 101 -> 35 t/s (one 10,240-token prefill took 293 s), whileq4_0gives 1726 t/s. Thecommon_fit_params: failed to fit params to free device memorywarning shows up in both states, so it is not the signal β KV bytes are. VRAM in the q4_0 state is 15.7/16.0 GiB. - Prefix reuse works within a conversation: re-sending an identical 12,485-token prompt evaluates only 4 tokens (0.28 s). Switching conversations re-prefills, since
-np 1keeps one slot.
Of the 12 prompts, one (a second Chinese prompt) stopped after 1 token with stop_type=eos and empty content in both arms β a raw-/completion-without-chat-template artifact, not a model defect. It is kept in the raw JSON and excluded from the statistics.
Not measured: the official 14-benchmark suite, repeats per prompt, contexts beyond 262,144, CPU-only runs, and the DFlash2 head.
Same data as a chart (baseline vs MTP throughput per prompt, and the acceptance rate that explains it):
The pattern to read off it: the speedup tracks acceptance, and acceptance tracks how predictable the continuation is β Python/list/JSON continuations land at 77-94% acceptance and 1.5-1.7x, free-form reasoning prose at 48-58% and 1.25-1.3x.
Code and method now live on GitHub β https://github.com/zhaoyilun/bonsai2-27b-mtp-repro β with the build script (SM89), the A/B harness, the launch units and the full write-up. This dataset stays the raw-data half of the same work, and the two link to each other.
Also added long-context numbers (chat path, natural text, depths exact via POST /tokenize, -c 262144 + KV q4_0):
| depth | prefill t/s | decode t/s (MTP) | acceptance | decode t/s (no spec) | speedup |
|---|---|---|---|---|---|
| 8,020 | 1855 | 86.9 | 65.8% | 63.0 | 1.38x |
| 31,939 | 1674 | 62.4 | 52.2% | 53.7 | 1.16x |
| 64,084 | 1332 | 55.4 | 67.9% | 43.4 | 1.28x |
| 127,870 | 974 | 46.1 | 82.3% | 33.3 | 1.38x |
| 191,099 | 722 | 35.0 | 84.1% | 26.5 | 1.32x |
Decode falls hard with depth (KV traffic), but MTP's edge does not decay β acceptance rises from 66% to 84%, so speculation buys more exactly where the target is slowest. One measurement gotcha worth knowing: timings.prompt_n reports only the newly evaluated tokens (a 191k-token prompt can come back as prompt_n = 63745) because llama-server reuses cached context checkpoints β the log shows n_tokens = 191139, truncated = 0, so nothing is dropped. A short "cache buster" request does not evict them; use /tokenize for the real length.
Filed the fix upstream as a pull request: https://github.com/PrismML-Eng/llama.cpp/pull/205 β qwen35: apply the Hadamard inverse to the MTP token-embedding lookup (1 file, +14 lines; it mirrors what the trunk already does in llm_graph_context::build_inp_embd, keyed on the table actually looked up).
Two things checked while preparing it:
- The bug is not version-specific. I downloaded the newest official release
prism-b10709-9a9394a(2026-09-18, ~26 builds newer than the b10683 I started from) and it still refuses the MTP context β with an extra diagnostic line now (latent lookup 'mtp_tok_embd-64' consumed by op=RMS_NORM name='norm-64'). - The fix is semantically exact. Built from the patched source for sm_89, draft acceptance on three probes is byte-for-byte the same as the community-patched build: 49/90, 59/71, 57/75 (54.4% / 83.1% / 76.0%).
Credit for finding it goes to runtime/bonsai-mtp-embedding.patch in ProCreations/Ternary-Bonsai-2-27B-MTP; the PR is the equivalent against current prism, so nothing here is claimed as new beyond rebasing it onto today's guard.
Also left supporting measurements on two related open issues from the same run: #203 (ngram-* silent no-op β independent CUDA corroboration) and #85 (q4_0 K-cache at 262144 on a 16 GB card without --kv-mean-center, plus the q8_0 failure mode).
