Ternary Bonsai 2 27B, NInfer artifact for the RTX 4090

Prism ML's Ternary Bonsai 2 27B converted to NInfer's .ninfer format, with the vision tower and a Qwen3.8 MTP head for speculative decoding. One self-contained file with text, vision and MTP.

These files only run on the NInfer build from JGamboa/ninfer-4090-windows (branch main), on an NVIDIA RTX 4090 (sm_89), on Windows or Linux. They do not load in llama.cpp, vLLM or Transformers.

Files

File MTP layer Size Use
bonsai2_27b_vl_mtp_q4q5.ninfer Q4/Q5 mix 6.4 GiB Recommended. 3.5–4.8 % faster decode than Q8, same acceptance
bonsai2_27b_vl.ninfer Q8 6.56 GiB Previous release, kept for reference

Each file has a .conversion.json report next to it: sources, formats and the conversion command.

Performance

Measured on one RTX 4090 at stock clocks, Windows 11, CUDA 13.4, with the file bonsai2_27b_vl_mtp_q4q5.ninfer. Current figures (2026-09-27, the 4090 without a display):

Measurement Result
Decode, MTP 2, six mixed prompts (story, code, math, Spanish, lists), greedy, thinking off 218 tok/s
Decode, no speculation (tg128) 129 tok/s
Prefill pp512 / pp2048 (ninfer_bench, int8 KV) 5,800-5,890 / 6,110 tok/s
Prefill, 8K / 64K / 128K-token prompt (needle test, answers exact) 1.3 s / 14.4 s / 37.4 s

Decode by workload (2026-09-24/25, when the card also drove a 4K desktop at 60 Hz, which cost about 15 % of decode), tok/s:

Workload MTP 2 MTP 2 + n-gram (--ngram chain)
Prose, greedy, thinking off 167 167
Edit-style agent prompts (return a file with a small change), greedy, thinking off 250 532
Through ninfer-serve, edit of a 7K–11K-token file, thinking on 215 378
Three concurrent requests, aggregate (int8 KV) 360 —
Against Prism's llama.cpp fork (b10709, PTQ1_0) on the same card, 2026-09-24 NInfer Prism fork
Decode, no speculation (tg128), 4K desktop on the 4090 in both 101 tok/s 77 tok/s
Prefill (pp512) 4,182-4,251 tok/s (5,800-5,890 now) 1,363 tok/s
Perplexity, wikitext / code 8.087 / 1.895 8.178 / 1.899
  • Quality: 43 of 45 deterministic tasks (code, JSON, tool calling, math, long context, Spanish), against 44 for Qwen3.8-27B.
  • Weights in VRAM: 6.11 GiB, or 6.39 GiB with vision.

Every figure, with its method, is in the repository's design notes (section 9.1).

Usage

Build NInfer from the GitHub branch above (the README has a Windows quick start), then run:

ninfer-serve bonsai2_27b_vl_mtp_q4q5.ninfer --host 127.0.0.1 --port 8080 ^
  --max-context 262144 --kv-capacity auto --kv-dtype rk4v4-e8 --max-concurrency 3 ^
  --spec mtp --draft-tokens 2 --lm-head-draft --ngram chain --vision
  • The server exposes an OpenAI- and Anthropic-compatible API at http://localhost:8080/v1, and a live monitor at /monitor.
  • On a 24 GB card this configuration gives each request the full 262K context. The three lanes share a pool of about 767K KV tokens.
  • --ngram chain extends each MTP round with drafts copied from the context. It helps on code editing, JSON, tool calls and documents, and does nothing on free prose.
  • The model thinks by default. Prism recommends temperature 1.0, top_p 0.95, top_k 20 in thinking mode.

What's inside

  • Ternary weights: Prism's codes are kept bit-exact, repacked as t5_g128_fp16 (5 weights per byte in base 3, one FP16 scale per 128 weights).
  • Vision tower: from Prism's Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf.
  • MTP head, tokenizer, chat template and configuration: from the official Qwen3.8-27B NInfer artifact. Bonsai keeps the Qwen3.8 architecture and tokenizer. In the recommended file, the MTP layer is requantized to Q4/Q5.

Modifications

This artifact is a derivative of Prism ML's Ternary Bonsai 2 27B. Changes made:

  • The ternary weights are repacked from Prism's PTQ1_0 GGUF packing to t5_g128_fp16. The codes are kept bit-exact and the FP16 group scales preserved.
  • The vision tower is converted from Prism's mmproj GGUF (Q8_0 dequantized to the NInfer layout).
  • The MTP speculative-decoding head, tokenizer, chat template and configuration are added from the Qwen3.8-27B NInfer artifact (derived from Qwen/Qwen3.8-27B). In bonsai2_27b_vl_mtp_q4q5.ninfer, the MTP layer is requantized from Q8 to a Q4/Q5 mix.
  • Everything is packaged as a single NInfer v3 .ninfer artifact.

License

Apache-2.0, the same license as all three sources: Prism ML's Ternary Bonsai 2 27B, Qwen/Qwen3.8-27B, and neroued/Qwen3.8-27B-NInfer.

Credits

  • Model: Prism ML (Ternary Bonsai 2 27B), based on Qwen3.8-27B by the Qwen team.
  • Engine: NInfer by Neroued. RTX 3090 and 4090 ports by Don-Chad and sergiuszm; E8 KV by UDPSendToFailed.
  • Windows port, Bonsai support, kernels, speculation and conversion: JGamboa, with Claude Code.
Downloads last month
13,125
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jgamboa/Ternary-Bonsai-2-27B-NInfer-4090

Base model

Qwen/Qwen3.8-27B
Quantized
(34)
this model

Space using jgamboa/Ternary-Bonsai-2-27B-NInfer-4090 1