Instructions to use prism-ml/Ternary-Bonsai-2-27B-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- LM Studio
- Jan
- vLLM
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prism-ml/Ternary-Bonsai-2-27B-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prism-ml/Ternary-Bonsai-2-27B-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Ollama
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Ollama:
ollama run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Unsloth Desktop
- Pi
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Docker Model Runner:
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Lemonade
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run and chat with the model
lemonade run user.Ternary-Bonsai-2-27B-gguf-F16
List all available models
lemonade list
- Hermes Agent
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
PLEASE stop lying about the "intelligence" of the model.
Hi. Please stop claiming that it's "as good as FP16" because it isn't.
Saying that is "98.2% as intelligent" to real humans it means that it should be able to do what the FP16 model does, BUT IT DOESN'T.
Ternary itself is awesome, and the fact that binary works at all is even more so.
But whoever at Bonsai decided to lie with clickbait tittles and claims should be ashamed of themselves.
Y'all are placing a stain on these kind of quants, instead of encouraging others to give it a try, y'all are placing a bad reputation on this side of the research.
I think the frustration behind this post points to a real measurement problem, although I would frame it a little differently.
I don't think the 98.2% number itself is fake.
According to Prism's published results, Bonsai 2 scores 84.78 vs 86.32 for FP16 across the 14 benchmarks currently shown on the model card. The broader 20-benchmark results similarly report 83.9 vs 85.4. For a ~5.9 GB representation of a 27B model, that is an extremely impressive result.
But I think we are starting to collect enough real-world reports to say that:
98.2% benchmark-score retention is not necessarily 98.2% behavioral-equivalence retention.
And that distinction may explain why the benchmark numbers and people's subjective experience seem so far apart.
The biggest clue is already present in the published results.
On relatively bounded tasks, Bonsai 2 looks remarkably intact. Math is almost unchanged. Short-form coding benchmarks are almost unchanged. Instruction following is strong. BFCL is still reasonably close.
But the long-horizon results are very different.
Terminal-Bench 2.1 is reported at 52.8 vs 69.7 for FP16, and SWE-bench Verified at 60.8 vs 80.6. In other words, on these long sequential tasks, retention is closer to roughly three quarters than 98%.
That is a very different failure mode from simply making every answer 2% worse.
And interestingly, the community reports are beginning to show the same pattern.
Discussion #18 contains a particularly useful example: Bonsai 2 consumed the full 32,768 reasoning-token budget without producing the requested final code. More importantly, the trace did not merely "need more thinking." It repeatedly returned to the same ideas, confused conventions, corrected errors and then reintroduced those same errors.
That looks less like a simple knowledge deficit and more like a convergence / self-correction stability problem.
Discussion #40 provides another interesting data point. One user saw a basic tool command fail repeatedly and degenerate into repetition. But another user in the same thread said the model can complete real projects — it just requires approximately 30% more reasoning/turns than their IQ2_S model to get things right.
That second observation may be especially important.
Suppose FP16 and Bonsai both eventually solve a task.
A conventional benchmark may record:
FP16: correct
Bonsai: correct
and therefore see almost no difference.
But imagine the actual trajectories were:
FP16: 8k generated tokens, 6 tool calls, 0 retries
Bonsai: 18k generated tokens, 11 tool calls, 3 corrections, 1 failed branch, correct answer
From an accuracy benchmark those results may be equivalent.
From the perspective of an agent harness, they are very different models.
This would also explain something that initially looks contradictory: coding benchmarks can remain almost unchanged while people report that coding agents feel substantially worse.
HumanEval, MBPP, LiveCodeBench and similar tests mostly ask whether the model can eventually construct the correct local solution.
An autonomous coding agent has a different problem.
It must repeatedly:
- decide what to do next,
- choose the correct tool,
- generate a syntactically valid tool call,
- interpret the result,
- update its internal model of the problem,
- remember previous decisions,
- avoid undoing already-correct work,
- recognize when a hypothesis failed,
- choose a new branch,
- and eventually decide that it is finished.
A tiny degradation in each individual decision can compound across dozens of steps.
If the probability of making the correct next decision changes only slightly, a one-shot benchmark may barely notice it. A 30-step trajectory can amplify it enormously.
That could be the missing dimension here.
There is another clue: reasoning-token efficiency.
Several reports mention unusually long reasoning, repetition or loops. Discussion #31 is especially interesting because xhigh reportedly failed to finish the evaluation due to a thinking loop, while switching to medium allowed the test to complete normally.
Discussion #35 contains an uncontrolled but striking comparison where the ternary model reportedly spent around 110k thinking tokens on a task for which a heavier quant of the original model used around 21k.
These anecdotes are not enough to establish an exact quantitative loss. xhigh itself can loop even on the full model, sampling configuration matters, and chat templates/runtime differences are significant confounders.
But they point toward something current benchmark tables largely do not report:
How much test-time compute is required to recover the benchmark score?
If aggressive ternarization reduces the confidence margin between competing decisions, additional reasoning may compensate for it surprisingly well.
That would actually explain BOTH observations:
The benchmark says the model retained almost all capability.
The human says the model became noticeably "dumber."
Both can be true.
The underlying knowledge and problem-solving machinery may still be there, allowing sufficiently long inference to eventually recover a correct result, while the model has become less decisive, less stable, more repetitive, and more expensive in generated tokens.
An agent harness can hide a surprising amount of this degradation.
A good harness can validate tool calls, reject malformed output, retry failed operations, externalize state, summarize history, constrain output with schemas, split large tasks into smaller tasks, keep a separate plan, prune bad trajectories, and even have multiple agents check one another.
Discussion #43 is a useful counterexample for exactly this reason: someone reports running a multi-agent orchestrator successfully with no failed tool calls. Discussion #36 also reports substantially better behavior with a corrected chat template and medium reasoning.
So I don't think it is accurate to say simply:
"Bonsai 2 cannot do agentic coding."
It clearly can.
The more interesting question is:
How much harness work and additional inference does it require to achieve the same task-level reliability as FP16 or a good 3–4 bit quant?
That is a measurable question.
I would love to see future Bonsai evaluations add metrics such as:
- success rate at fixed generated-token budgets (4k / 8k / 16k / 32k), rather than effectively allowing reasoning effort to compensate without accounting for its cost;
- median and p95 generated tokens per successful task;
- turns/tool calls/retries per successful task;
- malformed or semantically incorrect tool-call rate;
- recovery rate after an incorrect tool result or failed hypothesis;
- loop/non-termination/output-limit rate;
- frequency of reintroducing previously corrected errors;
- task success as context grows or as the number of interaction turns increases;
- identical FP16/Q4/IQ3/IQ2/Bonsai comparisons using the same harness, template, KV precision, sampler and reasoning budget;
- and perhaps most importantly, a tokens-to-correct-solution or success-per-1k-generated-tokens curve.
I suspect such measurements would reconcile a lot of the apparently contradictory reports.
We already have examples showing the other side too. Discussion #38 showed extremely good supplied-context reasoning, including updates, exceptions and cross-section reasoning. Discussion #36 reports good practical performance after template/settings changes. Discussion #43 demonstrates successful multi-agent use.
So this does not look like a model that has simply lost 20–30% of all intelligence.
Instead, my current hypothesis from the available evidence is that the degradation is highly non-uniform.
High-confidence, locally bounded computation seems to survive ternarization remarkably well.
What appears more fragile is long-horizon trajectory control: maintaining state, making repeated low-margin decisions, recovering from mistakes, avoiding revisiting discarded hypotheses, and knowing when to stop.
Those weaknesses are exactly the ones that conventional short-answer benchmarks can underweight — and exactly the ones humans notice very quickly when using an agent interactively.
This is also why I think this model is scientifically more interesting than the argument over the marketing wording.
If Bonsai really preserves most endpoint accuracy at ~1.7 bpw while significantly changing trajectory stability and inference efficiency, that tells us something important about what information aggressive quantization preserves and what it destroys.
Maybe the next useful question is no longer:
"How many percent of benchmark intelligence survived?"
but:
"How much inference effort is required to reconstruct the same useful behavior?"
That would distinguish weight-memory compression from effective system-level capability compression.
Personally, I would describe the current evidence as:
The memory compression result is extraordinary. Short and bounded capabilities appear surprisingly well preserved. But "98.2% of FP16 benchmark score" should not yet be interpreted as "98.2% behavioral equivalence to FP16," especially for long-horizon agentic workloads.
That is not a reason to dismiss ternary models.
Quite the opposite.
I think quantifying this gap — trajectory stability, context degradation, retries, convergence and token cost — may be one of the most interesting things the community can investigate next.
This is a classic example of an org picking the best number that they have and only showing that number.
I think this is a huge mistake, as you get people disappointed when it fails to live up to expectations, rather than people finding the model useful for what it's good at.
The facts are these:
This is likely the best sub-Q2 27B model you can run by a pretty wide margin.
It's very likely the best < 6gb model for vision and chat you can run, also by a pretty wide margin. It handily beats the best 9b models in most every category.
It is NOT trustworthy for multi-turn agentic runs, full stop. The benchmark they don't advertise is that it is something like 10 points lower than the Q4 quant of the model for multi-turn agentic performance, and in real world scenarios at long contexts it loses track of instructions and becomes unpredictable.
It is NOT efficient with thinking tokens. Real world testing shows that to get the best results you need xHigh reasoning, and it will use 2-3x the tokens of a Q4 model. If you're trying to use the small size for increased speed, this completely defeats that purpose.
If your only limitation is VRAM and you have sufficient RAM to run a ~17gb model, Qwen 3.6 35b A3b is going to be similar speed and much more reliable for agentic work, while performing as well or better on every other category.
I tested the prompt:
Generate an SVG of a pelican riding a bicycle
This produced a very interesting failure mode.
The model did not appear to have much difficulty with the overall composition itself. Fairly early in the reasoning process, it had already established most of the important elements:
- the bicycle geometry
- the pelican body, head, beak, pouch, wing, legs and feet
- pedal positions
- handlebars and wing placement
- background elements such as the road, clouds and sun
From those intermediate coordinates, the image was already quite recognizable as a pelican riding a bicycle.
What happened afterwards was much more interesting: the model kept refining the geometry for hours.
It repeatedly checked very small overlaps and clearances such as:
- leg vs. rear wheel
- leg vs. top tube
- wing vs. beak / pouch
- foot vs. pedal
- pouch vs. beak
- handlebar vs. wingtip
- draw order for overlapping elements
Some of these checks were literally at the 1–10 px level, and several previously checked relationships were checked again later.
Eventually, after roughly 52k reasoning tokens and more than 6 hours, the model finally finished reasoning and began writing the actual SVG answer.
Raw completed reasoning log:
pelican_raw_reasoning-bonsai2-RTX2060super-end.txt
SVG reconstructed from the final SVG data produced inside the completed reasoning:
pelican_bonsai2_final_reasoning_decoded.svg
The decoded SVG above is based on the SVG data produced at the end of the reasoning trace.
I only removed reasoning-only labels/text that were mixed into the draft SVG and applied the model's own final pouch-coordinate correction stated immediately afterwards, so that the result could be rendered as valid SVG.
So in this case, the failure mode looks much less like:
- “the model cannot construct the SVG”
and much more like:
- “the model has difficulty deciding that the SVG is already good enough and stopping.”
From a human point of view, the composition was already successful much earlier.
The vast majority of the later reasoning was spent trying to eliminate increasingly small geometric imperfections rather than fixing a fundamentally broken scene.
Examples of the geometry it kept checking
Bicycle
- Rear wheel center:
(235,355) - Front wheel center:
(468,355) - Wheel radius:
85 - Bottom bracket:
(350,336)
Frame:
- Top tube:
(292,262) -> (432,256) - Seat tube:
(292,258) -> (350,336) - Down tube:
(432,256) -> (350,336) - Chain stay:
(235,355) -> (350,336) - Seat stay:
(292,268) -> (235,355) - Front/head tube:
(432,256) -> (468,336)
Final crank/pedal revision:
- Pedal A:
(330,365) - Pedal B:
(370,307)
Pelican
Body:
ellipse cx=330 cy=205 rx=70 ry=55
Head:
circle cx=396 cy=148 r=34
Shoulder / neck patch:
ellipse cx=370 cy=172 rx=22 ry=20
Wing:
M282 198
C318 158 380 146 422 162
C452 172 466 194 466 214
C466 226 454 230 444 224
C414 204 380 196 344 194
C312 192 288 198 282 198 Z
Beak:
M425 140
C500 114 549 130 549 147
C549 161 506 174 427 162
C425 156 425 146 425 140 Z
Final pouch revision:
M446 158
C458 194 508 200 532 164
C500 165 476 159 446 158 Z
Final leg revision:
- Left leg:
(345,252) -> (330,358) - Right leg:
(364,248) -> (370,307)
A representative example
At one point, the model noticed that the left leg intersected the outer radius of the rear tire.
It then changed the crank/pedal geometry to move the leg outside the wheel.
Immediately afterwards, it discovered that the revised leg almost exactly crossed the bicycle's top tube around (345,260).
Instead of changing it again, it finally accepted that overlap because the leg was drawn after the frame:
Leg drawn after the frame → no problem ✓.
This kind of local collision analysis happened repeatedly throughout the trace.
My impression is that this may be a useful example of a stopping-criterion / over-refinement problem in long reasoning.
The model had already solved the semantic task — the scene clearly read as a pelican riding a bicycle — but it continued optimizing local geometry long after a human would probably have accepted the result.
Follow-up / addendum:
The run eventually did finish and produced a final SVG.
What makes this case especially interesting is that the model had already formed a recognizable scene much earlier, but continued spending a very long time on repeated micro-checks of geometry and overlap before committing to the final answer.
Artifacts:
Actual final SVG emitted by the model
diagram2.svgScreenshot of the final result in the UI
Screenshot 2026-09-23 at 20-33-27Completed raw reasoning log
pelican_raw_reasoning-bonsai2-RTX2060super-end.txtllama-server log
log-llama-server260923.txt
A few updated observations:
This was not simply a case of “the model cannot generate the SVG.”
The final result is a valid and recognizable cartoon pelican riding a bicycle.The failure mode is much closer to:
the model has difficulty deciding that the result is already good enough and stopping.During reasoning, it repeatedly revisited very small geometry checks, such as:
- leg vs. rear wheel
- leg vs. top tube
- pouch vs. beak
- wing vs. handlebar / head
- foot vs. pedal
- draw-order based overlap hiding
The visible reasoning run itself already reached roughly 52k+ reasoning tokens over about 6 hours, and the complete end-to-end wall-clock time was over ~11 hours when including prefill and other intermediate interactions (for example, time/date/tool-launch related questions before the final SVG was emitted).
So my impression remains the same, but the completed run strengthens it:
- The model can construct the scene.
- The main issue is not scene formation.
- The main issue is over-refinement / stopping criterion in long reasoning.
In other words, the semantic task was solved much earlier, but the model kept spending a very large amount of additional reasoning on eliminating tiny local imperfections before finally writing out the SVG.
That makes this a useful example of a long-reasoning failure mode where the system does eventually succeed, but only after an impractically long period of repeated fine-grained self-checking.
I had about a 1.5h conversation on the same topic, that models typically aren't allowed to talk, with both this PTQ1_0 and another Q2_K_XL, that has 65% bigger filesize. And I felt like the latter was a very knowledgeable kid. It gave good answers, but was very easy to trick and lead to where I want to. On the other hand this model was much more aware, it was harder to cheat and closer to the end of conversation it outright said that it sees that I try to make, gave me 8 examples of my actions starting from ~30k earlier context, predicted what I try to achieve and was ready only for some compromise. It felt more reasonable, despite being just 5.5gb model and sub 2bit weights. In other words not just knowledge, but also logic.
It's not a benchmark and just one sample, but it's also not a benchmark, but actual usage and 1.5h conversation with ~32k tokens probably isn't such a bad indicator of abilities. For it's weight it felt very strong (but also it was spending tokens on thinking like there is no tomorrow, looking for many different directions and factor that can be useful, so maybe it also helps).