Instructions to use prism-ml/Ternary-Bonsai-2-27B-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- LM Studio
- Jan
- vLLM
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prism-ml/Ternary-Bonsai-2-27B-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prism-ml/Ternary-Bonsai-2-27B-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Ollama
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Ollama:
ollama run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Unsloth Desktop
- Pi
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Docker Model Runner:
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Lemonade
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run and chat with the model
lemonade run user.Ternary-Bonsai-2-27B-gguf-F16
List all available models
lemonade list
- Hermes Agent
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
for 16gb vram context size <=114688 is fast. Pelican is gorgeous :)
was curious why model speed was 1t/s on 262144.. so 114688 is last fast size for 16gb vram, giving 29.5 t/s .
./llama-cli -m Ternary-Bonsai-2-27B-PQ2_0.gguf --temp 1 --top-p 0.95 --min-p 0 --top-k 20 --repeat-penalty 1 --presence-penalty 0 --seed 42 --ctx-size 114688
Generate a svg of a pelican riding a bicycle
build : b10709-9a9394a89
<svg viewBox="0 0 900 600" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="A white pelican with an orange beak riding a teal bicycle, scarf flapping, speeding off across a road">
<defs>
<linearGradient id="skyGrad" x1="0" y1="0" x2="0" y2="1">
<stop offset="0" stop-color="#86d2f2"/><stop offset="1" stop-color="#e8f8fd"/>
</linearGradient>
<linearGradient id="beakGrad" x1="0" y1="0" x2="0" y2="1">
<stop offset="0" stop-color="#ffbf6f"/><stop offset="1" stop-color="#ff8a3c"/>
</linearGradient>
<linearGradient id="bodyGrad" x1="0" y1="0" x2="0" y2="1">
<stop offset="0" stop-color="#ffffff"/><stop offset="1" stop-color="#e9eef7"/>
</linearGradient>
</defs>
<title>Pelican riding a bicycle</title>
<rect width="900" height="600" fill="url(#skyGrad)"/>
<!-- sun -->
<g stroke="#ffd98f" stroke-width="5" stroke-linecap="round" fill="none">
<path d="M60 84 H176"/><path d="M118 26 V142"/>
<path d="M79 45 L157 123"/><path d="M157 45 L79 123"/>
</g>
<circle cx="118" cy="84" r="40" fill="#ffd36b"/>
<!-- clouds -->
<g fill="#ffffff">
<ellipse cx="240" cy="92" rx="26" ry="20" opacity=".95"/><ellipse cx="274" cy="80" rx="32" ry="24" opacity=".95"/><ellipse cx="306" cy="92" rx="24" ry="18" opacity=".95"/>
<ellipse cx="684" cy="64" rx="22" ry="17" opacity=".85"/><ellipse cx="712" cy="54" rx="28" ry="21" opacity=".85"/><ellipse cx="738" cy="64" rx="20" ry="15" opacity=".85"/>
<ellipse cx="778" cy="104" rx="12" ry="10" opacity=".75"/><ellipse cx="798" cy="92" rx="15" ry="12" opacity=".75"/><ellipse cx="816" cy="104" rx="11" ry="9" opacity=".75"/>
</g>
<path d="M766 150 Q774 138 784 146 Q794 138 802 150" stroke="#7091ab" stroke-width="3" fill="none" stroke-linecap="round"/>
<!-- hills -->
<path d="M0 508 C120 478 250 496 370 484 C500 474 640 488 770 480 C840 476 880 488 900 494 V508 Z" fill="#cfe6c8"/>
<path d="M900 508 C812 474 720 488 628 480 C560 475 508 486 470 508 Z" fill="#b9d4b4"/>
<!-- road -->
<rect x="0" y="508" width="900" height="92" fill="#a9a499"/>
<rect x="0" y="508" width="900" height="5" fill="#9a958a"/>
<g fill="#969287" opacity=".7">
<circle cx="70" cy="548" r="2"/><circle cx="150" cy="590" r="2.2"/><circle cx="300" cy="585" r="2"/><circle cx="86" cy="572" r="2"/>
</g>
<path d="M0 562 H900" stroke="#f1eee0" stroke-width="7" stroke-dasharray="46 30" opacity=".85"/>
<g stroke="#e9e5d5" stroke-width="6" stroke-linecap="round" opacity=".8">
<path d="M64 540 H132"/><path d="M142 580 H200"/>
</g>
<ellipse cx="430" cy="517" rx="250" ry="11" fill="#4a4a45" opacity=".14"/>
<!-- speed lines -->
<g stroke="#6fbfe5" stroke-width="5" stroke-linecap="round" fill="none" opacity=".6">
<path d="M18 150 H98"/><path d="M38 198 H122"/><path d="M8 248 H86"/>
<path d="M28 300 H108"/><path d="M52 354 H132"/><path d="M66 408 H148"/>
</g>
<!-- ===== BICYCLE ===== -->
<!-- rear wheel -->
<circle cx="322" cy="436" r="66" stroke="#26292f" stroke-width="11" fill="none"/>
<circle cx="322" cy="436" r="61" stroke="#d0d5de" stroke-width="5" fill="none"/>
<g stroke="#c9d0dc" stroke-width="2.5" fill="none">
<path d="M334 436 H378"/><path d="M332 442 L370 464"/><path d="M328 446 L349 484"/><path d="M322 448 V492"/>
<path d="M316 446 L295 484"/><path d="M312 442 L273 464"/><path d="M310 436 H266"/><path d="M312 430 L273 394"/>
<path d="M316 430 L295 388"/><path d="M322 424 V380"/><path d="M328 426 L349 382"/><path d="M332 430 L370 394"/>
</g>
<circle cx="322" cy="436" r="10" fill="#5e6673"/><circle cx="322" cy="436" r="4" fill="#30353d"/>
<!-- fenders -->
<path d="M258 392 A78 78 0 0 1 386 392" stroke="#166069" stroke-width="8" fill="none" stroke-linecap="round"/>
<path d="M490 395 A74 74 0 0 1 614 395" stroke="#166069" stroke-width="8" fill="none" stroke-linecap="round"/>
<!-- front wheel -->
<circle cx="552" cy="436" r="66" stroke="#26292f" stroke-width="11" fill="none"/>
<circle cx="552" cy="436" r="61" stroke="#d0d5de" stroke-width="5" fill="none"/>
<g stroke="#c9d0dc" stroke-width="2.5" fill="none">
<path d="M564 436 H608"/><path d="M562 442 L600 464"/><path d="M558 446 L579 484"/><path d="M552 448 V492"/>
<path d="M546 446 L525 484"/><path d="M542 442 L503 464"/><path d="M540 436 H496"/><path d="M542 430 L503 394"/>
<path d="M546 430 L525 388"/><path d="M552 424 V380"/><path d="M558 426 L579 382"/><path d="M562 430 L600 394"/>
</g>
<circle cx="552" cy="436" r="10" fill="#5e6673"/><circle cx="552" cy="436" r="4" fill="#30353d"/>
<!-- frame -->
<g stroke="#1f737b" fill="none" stroke-linecap="round">
<path d="M332 432 L354 232" stroke-width="13"/>
<path d="M358 250 L500 318" stroke-width="13"/>
<path d="M360 404 L502 328" stroke-width="13"/>
<path d="M502 318 L518 386" stroke-width="13"/>
<path d="M500 318 L552 302" stroke-width="9"/>
</g>
<!-- fork -->
<g stroke="#1f737b" fill="none" stroke-linecap="round" stroke-width="8">
<path d="M516 390 H532"/><path d="M516 390 L550 432"/><path d="M532 390 L556 432"/>
</g>
<!-- handlebar -->
<path d="M550 302 C588 290 620 294 634 318" stroke="#1f737b" stroke-width="9" fill="none" stroke-linecap="round"/>
<!-- saddle -->
<path d="M354 232 L348 208" stroke="#4e555f" stroke-width="9" fill="none" stroke-linecap="round"/>
<path d="M332 202 Q354 194 378 202 Q380 210 370 214 Q354 216 342 212 Z" fill="#7a4a2b"/>
<!-- chainring, chain, pedals -->
<circle cx="322" cy="436" r="23" stroke="#3c414c" stroke-width="5" fill="none"/>
<circle cx="322" cy="436" r="4" fill="#3c414c"/>
<g stroke="#3c414c" stroke-width="3" stroke-linecap="round">
<path d="M345 413 H394"/><path d="M345 459 H394"/>
</g>
<g stroke="#444a55" stroke-width="8" stroke-linecap="round" fill="none">
<path d="M322 436 L286 500"/><path d="M322 436 L358 372"/>
</g>
<g stroke="#2e333c" stroke-width="9" stroke-linecap="round">
<path d="M268 504 L304 496"/><path d="M342 363 L374 381"/>
</g>
<!-- pelican legs -->
<path d="M396 262 C388 320 368 348 356 368" stroke="#ff9638" stroke-width="10" fill="none" stroke-linecap="round"/>
<path d="M356 252 C330 338 306 420 298 488" stroke="#ff9638" stroke-width="11" fill="none" stroke-linecap="round"/>
<path d="M272 498 L306 488" stroke="#ff8f3d" stroke-width="12" stroke-linecap="round"/>
<path d="M346 368 L372 378" stroke="#ff8f3d" stroke-width="11" stroke-linecap="round"/>
<!-- ===== PELICAN ===== -->
<!-- tail -->
<g stroke="#93a1b2" stroke-width="6" stroke-linecap="round" fill="none">
<path d="M340 240 L300 226"/><path d="M336 254 L294 252"/><path d="M340 268 L304 280"/>
</g>
<!-- body -->
<ellipse cx="404" cy="224" rx="78" ry="46" fill="url(#bodyGrad)"/>
<path d="M356 262 C398 280 452 276 482 246 C458 272 410 282 356 262 Z" fill="#e3eaf4"/>
<g stroke="#cfe0ee" stroke-width="3" fill="none" stroke-linecap="round">
<path d="M456 238 Q466 244 470 254"/><path d="M444 256 Q454 262 460 272"/>
</g>
<!-- folded back wing -->
<path d="M362 216 C334 202 306 200 282 212 C272 216 276 228 288 228 C312 238 340 240 366 236 Z" fill="#c9d4e2"/>
<g stroke="#a9b7ca" stroke-width="2.5" stroke-linecap="round" fill="none">
<path d="M312 214 L288 218"/><path d="M326 222 L302 228"/><path d="M342 230 L320 236"/>
</g>
<!-- neck -->
<path d="M450 214 Q484 184 502 128" stroke="#ffffff" stroke-width="38" fill="none" stroke-linecap="round"/>
<path d="M452 206 Q476 182 494 142" stroke="#e6edf7" stroke-width="9" fill="none" stroke-linecap="round"/>
<!-- beak (under head) -->
<path d="M528 92 C562 86 602 96 634 118 L642 126 C646 132 638 138 631 131 C620 137 613 148 615 162 C616 176 603 186 588 178 C578 172 568 158 562 148 C556 138 547 122 537 108 Z" fill="url(#beakGrad)"/>
<!-- head -->
<circle cx="512" cy="96" r="23" fill="#ffffff"/>
<path d="M496 84 C489 92 488 102 494 114 C497 106 499 98 495 90 Z" fill="#e3ebf5"/>
<path d="M498 82 C506 74 518 72 528 78 C520 78 510 81 502 86 Z" fill="#e9f0f8"/>
<circle cx="546" cy="97" r="1.8" fill="#c95d28" opacity=".8"/>
<path d="M536 104 C572 100 610 106 634 118" stroke="#c95d28" stroke-width="3" stroke-linecap="round" fill="none"/>
<path d="M600 148 C604 158 610 166 616 168 C610 166 602 156 598 144 Z" fill="#f98f4b" opacity=".65"/>
<circle cx="520" cy="88" r="4.3" fill="#2b3038"/>
<circle cx="522" cy="85.5" r="1.5" fill="#ffffff"/>
<!-- scarf -->
<path d="M436 192 C442 178 458 174 466 184 C470 192 460 202 448 200 Z" fill="#ff6b5c"/>
<path d="M438 198 C396 196 360 210 330 192 C318 185 312 196 324 202" stroke="#ff6b5c" stroke-width="7" stroke-linecap="round" fill="none"/>
<path d="M430 206 C396 208 368 218 344 206" stroke="#ff9a8c" stroke-width="4.5" stroke-linecap="round" fill="none"/>
<!-- wing gripping the handlebar -->
<path d="M446 206 C488 238 544 266 590 286 C608 294 620 304 610 314 C596 324 552 310 508 290 C474 274 454 240 446 206 Z" fill="#f4f7fc" stroke="#d4dde9" stroke-width="2"/>
<g stroke="#cdd6e2" stroke-width="3" stroke-linecap="round" fill="none">
<path d="M486 232 C530 250 570 268 594 280"/><path d="M496 252 C536 268 564 284 588 294"/>
</g>
<!-- air rushing past the beak -->
<g stroke="#b7e6f3" stroke-width="4" stroke-linecap="round" fill="none" opacity=".8">
<path d="M658 132 H692"/><path d="M666 152 H686"/>
</g>
</svg>
Nice! I'm the person who posted the RTX 2060 SUPER 8GB (Turing) test in discussion #47:
https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/discussions/47
You mentioned 16GB VRAM and that <=114688 context is fast. What exact GPU model / generation are you using?
I'm curious how much of the difference comes from VRAM capacity versus GPU generation and memory bandwidth.
I have RTX 5060 Ti 16gb vram. 32k context is not enough for first Pelican, as this llm thinks way more than original (reasoning text is 102k).
I'm a little jealous of that RTX 5060 Ti π
My RTX 2060 SUPER gets around 16.8 tok/s with PTQ1_0, so your ~29.5 tok/s result is about 1.76Γ faster.
What makes this especially interesting is that both cards have roughly the same 448 GB/s memory bandwidth on paper. General GPU compute benchmarks also seem to show around a 1.7β1.8Γ gap between the RTX 2060 SUPER and RTX 5060 Ti, which is surprisingly close to the difference we're seeing here.
So I suspect the newer Blackwell generation is helping quite a lot on the compute side β things like ternary decode, integer/DP4A operations, scheduling, and possibly Flash Attention β rather than this being explained by memory bandwidth alone.
A 27B ternary model doing nearly 30 tok/s on a 16GB card is really nice. I'm definitely a bit jealous of that 5060 Ti π
let me clarify for you: 29~32 t/s speed is on 1500mhz (62W), on base 2400mhz (113W) speed is 42t/s. 2100mhz (93W) 41t/s, 1800mhz(79W) 38t/s
if I minimise bash window: +2-3 t/s
New chat: Generate a svg of a pelican riding a bicycle
Seems absolutely worse
Ternary-Bonsai-2 27B UncensoredHereticPQ2_0 / RX 6950XT (16GB RDNA2 - the AMD answer to RTX 4000) / Windows / Vulkan
-ngl all -fa on -c 131072 --host 0.0.0.0 -t 16 --temperature 0.7 --dynatemp-range 0.15 --top-p 0.44 --top-k 12 --min-p 0.3 --repeat-penalty 1.12 --presence-penalty 1.12 -ctk q5_1 -ctv q5_1 --spec-type ngram-mod,ngram-map-k4v -np 1 --spec-draft-n-max 2 --spec-ngram-mod-n-match 20 --spec-ngram-mod-n-min 36 --spec-ngram-mod-n-max 68 --spec-ngram-map-k4v-size-n 10 --cache-ram 4096 --chat-template-file .\Qwen-3.8\Qwen-sharp-chat_template.jinja --reasoning-effort medium --perf --slot-save-path .\cache\ -b 8192 -ub 448 -cms 3172 -ctxcp 56 --lookup-cache-dynamic .\cache\n-gram.cache -lm dio
there was some effect of the ngram, here is the final stats:
[34m14.37.813.530[0m [32mI [0mslot print_timing: id 0 | task 0 | prompt eval time = 9977.33 ms / 2643 tokens ( 3.78 ms per token, 264.90 tokens per second)
[34m14.37.813.540[0m [32mI [0mslot print_timing: id 0 | task 0 | eval time = 814095.49 ms / 10702 tokens ( 76.08 ms per token, 13.14 tokens per second)
[34m14.37.813.544[0m [32mI [0mslot print_timing: id 0 | task 0 | total time = 824072.83 ms / 13345 tokens
[34m14.37.813.549[0m [32mI [0mslot print_timing: id 0 | task 0 | graphs reused = 8714
[34m14.37.813.854[0m [32mI [0mslot print_timing: id 0 | task 0 | draft acceptance = 0.19157 ( 1473 accepted / 7689 generated), mean len = 8.48
well, first - you are running modified model; second - lot of parameters are different from basic. so your work not eligible π
The 114,688 cutoff lines up almost exactly with the KV cache size. So it's VRAM capacity, not GPU generation.
From the PQ2_0 GGUF header: qwen35.block_count = 64, full_attention_interval = 4, head_count_kv = 4, key_length = value_length = 256. Only 16 layers keep a KV cache. The other 48 are Gated DeltaNet layers with a fixed-size state of about 150 MiB in total. So f16 KV per token is 16 Γ 4 Γ (256 + 256) Γ 2 B = 64 KiB.
- weights:
Ternary-Bonsai-2-27B-PQ2_0.ggufis 7,206,168,928 B = 6.71 GiB -c 114688: 114,688 Γ 64 KiB = 7.00 GiB of KV, so 6.71 + 7.00 + 0.15 = 13.86 GiB plus the compute buffer-c 131072: 8.00 GiB of KV, so 14.86 GiB plus the compute buffer-c 262144: 16.00 GiB of KV, so 22.86 GiB. The KV cache alone is bigger than the card.
llama.cpp's --fit is on by default and tries to leave 1 GiB free per GPU, so a 16 GB card has a budget of roughly 15 GiB. Your command sets -c but not -ngl. When the total doesn't fit, fit lowers -ngl and moves layers to the CPU, which is the drop to ~1 t/s at 262K. 114,688 is about the last size where everything still stays on the GPU. The offloaded N/M layers to GPU line in the log shows whether it happened.
To get more context on 16 GB, -ctk q8_0 -ctv q8_0 roughly halves the cache to ~34 KiB/token. 196,608 Γ 34 KiB = 6.38 GiB, so ~192K in q8_0 takes about the same room as 114K in f16. The full 262K in q8_0 is 8.5 GiB, which together with the weights is still a bit over the budget.

