Instructions to use jgamboa/Ternary-Bonsai-2-27B-NInfer-4090 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NInfer
How to use jgamboa/Ternary-Bonsai-2-27B-NInfer-4090 with NInfer:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Ternary Bonsai 2 27B, NInfer artifact for the RTX 4090
Prism ML's Ternary Bonsai 2 27B
converted to NInfer's .ninfer format, with the vision tower and a Qwen3.8 MTP head for
speculative decoding. One self-contained file with text, vision and MTP.
These files only run on the NInfer build from
JGamboa/ninfer-4090-windows
(branch main), on an NVIDIA RTX 4090 (sm_89), on Windows or Linux. They do
not load in llama.cpp, vLLM or Transformers.
Files
| File | MTP layer | Size | Use |
|---|---|---|---|
bonsai2_27b_vl_mtp_q4q5.ninfer |
Q4/Q5 mix | 6.4 GiB | Recommended. 3.5–4.8 % faster decode than Q8, same acceptance |
bonsai2_27b_vl.ninfer |
Q8 | 6.56 GiB | Previous release, kept for reference |
Each file has a .conversion.json report next to it: sources, formats and the conversion
command.
Performance
Measured on one RTX 4090 at stock clocks, Windows 11, CUDA 13.4, with the file
bonsai2_27b_vl_mtp_q4q5.ninfer. Current figures (2026-09-27, the 4090 without a display):
| Measurement | Result |
|---|---|
| Decode, MTP 2, six mixed prompts (story, code, math, Spanish, lists), greedy, thinking off | 218 tok/s |
Decode, no speculation (tg128) |
129 tok/s |
Prefill pp512 / pp2048 (ninfer_bench, int8 KV) |
5,800-5,890 / 6,110 tok/s |
| Prefill, 8K / 64K / 128K-token prompt (needle test, answers exact) | 1.3 s / 14.4 s / 37.4 s |
Decode by workload (2026-09-24/25, when the card also drove a 4K desktop at 60 Hz, which cost about 15 % of decode), tok/s:
| Workload | MTP 2 | MTP 2 + n-gram (--ngram chain) |
|---|---|---|
| Prose, greedy, thinking off | 167 | 167 |
| Edit-style agent prompts (return a file with a small change), greedy, thinking off | 250 | 532 |
Through ninfer-serve, edit of a 7K–11K-token file, thinking on |
215 | 378 |
| Three concurrent requests, aggregate (int8 KV) | 360 | — |
| Against Prism's llama.cpp fork (b10709, PTQ1_0) on the same card, 2026-09-24 | NInfer | Prism fork |
|---|---|---|
| Decode, no speculation (tg128), 4K desktop on the 4090 in both | 101 tok/s | 77 tok/s |
| Prefill (pp512) | 4,182-4,251 tok/s (5,800-5,890 now) | 1,363 tok/s |
| Perplexity, wikitext / code | 8.087 / 1.895 | 8.178 / 1.899 |
- Quality: 43 of 45 deterministic tasks (code, JSON, tool calling, math, long context, Spanish), against 44 for Qwen3.8-27B.
- Weights in VRAM: 6.11 GiB, or 6.39 GiB with vision.
Every figure, with its method, is in the repository's design notes (section 9.1).
Usage
Build NInfer from the GitHub branch above (the README has a Windows quick start), then run:
ninfer-serve bonsai2_27b_vl_mtp_q4q5.ninfer --host 127.0.0.1 --port 8080 ^
--max-context 262144 --kv-capacity auto --kv-dtype rk4v4-e8 --max-concurrency 3 ^
--spec mtp --draft-tokens 2 --lm-head-draft --ngram chain --vision
- The server exposes an OpenAI- and Anthropic-compatible API at
http://localhost:8080/v1, and a live monitor at/monitor. - On a 24 GB card this configuration gives each request the full 262K context. The three lanes share a pool of about 767K KV tokens.
--ngram chainextends each MTP round with drafts copied from the context. It helps on code editing, JSON, tool calls and documents, and does nothing on free prose.- The model thinks by default. Prism recommends
temperature 1.0, top_p 0.95, top_k 20in thinking mode.
What's inside
- Ternary weights: Prism's codes are kept bit-exact, repacked as
t5_g128_fp16(5 weights per byte in base 3, one FP16 scale per 128 weights). - Vision tower: from Prism's
Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf. - MTP head, tokenizer, chat template and configuration: from the official Qwen3.8-27B NInfer artifact. Bonsai keeps the Qwen3.8 architecture and tokenizer. In the recommended file, the MTP layer is requantized to Q4/Q5.
Modifications
This artifact is a derivative of Prism ML's Ternary Bonsai 2 27B. Changes made:
- The ternary weights are repacked from Prism's PTQ1_0 GGUF packing to
t5_g128_fp16. The codes are kept bit-exact and the FP16 group scales preserved. - The vision tower is converted from Prism's
mmprojGGUF (Q8_0 dequantized to the NInfer layout). - The MTP speculative-decoding head, tokenizer, chat template and configuration are added from
the Qwen3.8-27B NInfer artifact (derived from Qwen/Qwen3.8-27B). In
bonsai2_27b_vl_mtp_q4q5.ninfer, the MTP layer is requantized from Q8 to a Q4/Q5 mix. - Everything is packaged as a single NInfer v3
.ninferartifact.
License
Apache-2.0, the same license as all three sources: Prism ML's Ternary Bonsai 2 27B, Qwen/Qwen3.8-27B, and neroued/Qwen3.8-27B-NInfer.
Credits
- Downloads last month
- 13,125