fyb1214's picture
metadata: single direct parent in base_model (ukisai only); prism-ml shows via its own tree
f605547 verified
|
Raw History Blame Contribute Delete
8.94 kB
metadata
license: apache-2.0
library_name: ninfer
pipeline_tag: image-text-to-text
inference: false
base_model:
  - ukisai/Swift-Bonsai-2-GGUF
base_model_relation: quantized
language:
  - en
  - zh
tags:
  - ninfer
  - ternary
  - 1-bit
  - 2-bit
  - pq2
  - ptq1
  - bonsai
  - qwen3.8
  - hadamard
  - mtp
  - speculative-decoding
  - swift
  - multimodal

Swift-Bonsai-2 27B β€” NInfer

Swift-Bonsai-2 is UkisAI's reasoning-efficient derivative of Prism ML's Ternary Bonsai 2 27B, packaged here for the NInfer engine as PQ2_0_G128 and PTQ1_0_G128 artifacts.

The ternary codes were moved byte-for-byte out of the source GGUF into the NInfer container β€” nothing was dequantized and requantized, so there is no second quantization loss.

Quantizations

Quantization File Download size
1-bit / PTQ1_0 bonsai2_27b_swift_ptq1.ninfer 7.047 GB
2-bit / PQ2_0 bonsai2_27b_swift_pq2.ninfer 8.307 GB

Both files are in the root of this repository β€” download them from the Files tab. Each file is a complete container β€” text tower, vision tower, MTP head, proposal head, tokenizer, chat template and media-processor resources. No adapter file, patch, or extra flag is needed. No DFlash2 adapter.

Engine compatibility

This artifact uses the PQ2_0_G128 / PTQ1_0_G128 dialect and needs an engine from the Ambolio/ninfer-4090-windows lineage (v1.0.6 / v1.0.8).

Where to get the engine: the original Ambolio repos were taken down by their author (the account shows zero public repos; the HF mirror 404s). The lineage persists through forks β€” engine source, per-GPU build notes, and the converter toolchain that produced this artifact are at 5258MF/ninfer-swift-bonsai2-converter.

It is not interchangeable with the other Bonsai .ninfer artifacts on the Hub:

Artifact Ternary format Engine
this repo PQ2_0_G128 / PTQ1_0_G128 Ambolio lineage (v1.0.6 / v1.0.8)
WaveCut/Ternary-Bonsai-2-27B-NInfer-v3 t2_g128_fp16 iamwavecut/ninfer-all
neroued/Qwen3.8-27B-NInfer NVFP4 / groupwise-int Neroued/ninfer

To check which dialect a .ninfer uses, read the "format" fields in its first 1 MiB.

Requirements

The engine from the Ambolio lineage, built or released for your GPU. KV dtype is architecture-gated:

GPU KV types
RTX 30 series (sm_86) bf16, int8
RTX 40 series (sm_89) bf16, int8, fp8, rk4v4, rk4v4-e8
RTX 50 series (sm_120) bf16, int8, fp8, nvfp4, k8v4

Quick start

ninfer-serve.exe bonsai2_27b_swift_pq2.ninfer ^
  --host 127.0.0.1 --port 8087 ^
  --max-context 131072 --kv-capacity 131072 --kv-dtype int8 ^
  --spec mtp --draft-tokens 3

On a 12 GB card, --max-context 32768~131072 with --kv-dtype int8 was the configuration measured below. Scan the MTP window (1–4) on your own card β€” see Measured.

Measured

RTX 3060 12G, production parameters (ctx 65536 / int8 / MTP d4+lm / vision), greedy, 400 output tokens, 1 warmup + 2 runs (median), both models started alternately.

Config Swift (t/s) base (t/s)
d1 44.6 45.0
d2 52.5 51.5
d3 55.0 55.1
d4+lm 53.1 52.3

Perplexity, same corpus as the baseline: 8.073669 / 26.962015 / 148.3656 (Swift) versus 8.079208 / 26.972171 / 148.3863 (base) at window/stride 512/256, 32/16 and 8/4 β€” Swift is lower by 0.07 %, 0.04 % and 0.01 %. Swift and base stay within Β±2 % on throughput in every configuration.

MTP settings: the best draft depth depends on how predictable the content is. On the official bench (single-token free continuation) d4+lm leads: 76.6 vs 71.4 t/s at tg256, 62.0 vs 58.8 at tg128. Over HTTP with real 400-token generations d3 leads instead: 50.3 vs 47.6 t/s (greedy 52.0 vs 51.7). The gap on real content is ~1–5 %, near the noise floor; the bench gap is larger.

A caveat worth knowing: acceptance rate is a property of the content, not of the card. A verbatim-repeat prompt pushes acceptance past 96 % and makes the largest draft look best by construction. Measure with prompts that resemble your own workload.

Full throughput table (four prompt types Γ— four MTP tiers)

Format: t/s (MTP acceptance). Chinese / English / Code / Thinking.

Config Model Chinese English Code Thinking Mean
d1 Swift 42.2 (58%) 44.4 (68%) 46.0 (82%) 45.8 (84%) 44.6
d2 Swift 46.0 (46%) 49.5 (56%) 58.4 (80%) 55.9 (76%) 52.5
d3 Swift 44.0 (33%) 49.7 (44%) 64.2 (70%) 62.2 (68%) 55.0
d4+lm Swift 40.4 (28%) 48.0 (40%) 65.5 (67%) 58.6 (59%) 53.1
d1 base 43.1 (62%) 44.9 (71%) 46.6 (85%) 45.4 (83%) 45.0
d2 base 43.7 (42%) 46.9 (51%) 59.6 (82%) 55.7 (76%) 51.5
d3 base 43.4 (33%) 52.4 (48%) 64.9 (71%) 59.9 (65%) 55.1
d4+lm base 39.2 (27%) 48.2 (41%) 64.6 (66%) 57.1 (56%) 52.3

Absolute PPL values are corpus-dependent. Reference figures quoted elsewhere (6.448742 / 26.049634 / 121.157720) come from a different corpus and are not comparable to the numbers above; only the Swift-versus-base delta within one corpus is meaningful.

Scope of validation β€” what was and was not checked

Validated end-to-end by the publisher on one RTX 3060 12G (2026-09-28): server startup, coherent English and Chinese output, multi-turn prefix reuse (three same-prefix requests, 200/200/200/200), ten consecutive requests, sampling, all four reasoning tiers, vision (a red square, CAT, a blue 42), and tool calling. Prefix reuse measured 2848 ms β†’ 258 ms on a ~2.4k-token prompt.

On the advertised "~40 % fewer thinking tokens": that figure is the upstream GPQA-Diamond result. The same upstream table reports βˆ’7.1 % on C-Eval, +0.4 % on IFBench and +5.1 % on AIME 2025 β€” it is not a general property. Our AIME 2024 test (20 problems Γ— 2 rounds) points the same way: accuracy 37/40 vs 35/40, total output tokens βˆ’9 %, but the paired per-problem ratio has a median of 1.01 β€” for a typical problem the two models reason for the same length. GPQA was not tested here.

The prefix-reuse crash documented for older engine builds did not reproduce here.

Verify

Get-FileHash bonsai2_27b_swift_pq2.ninfer -Algorithm SHA256
# cc54be3800099ada67165ad352450be83d28e8c4d423be9cae573a7b6e6350a0

Get-FileHash bonsai2_27b_swift_ptq1.ninfer -Algorithm SHA256
# cc9e890728ea7357b1d8a0797a4accdc6ca143031471a63e49c9cf6c8314b6ae

SHA256SUMS and artifact-manifest.json carry the same values.

Reproducing it

python -u pack.py build bonsai2_27b_swift_pq2.ninfer \
  --gguf Swift-Bonsai-2-PQ2_0.gguf \
  --template qwen3_8_27b_huihui_abliterated.ninfer

Source and template hashes, the full procedure and the verification methodology are in REPRODUCE.md.

License and attribution

Bonsai 2 27B (original weights)   Β© Prism ML, Inc.     Apache-2.0
                                 huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf
Qwen3.8-27B (geometry base)       Β© Alibaba Cloud      Apache-2.0
Swift fine-tune (weight source)   ukisai               Apache-2.0
                                 huggingface.co/ukisai/Swift-Bonsai-2-GGUF
Packer                            pack.py β€” shensanshu/ninfer-ada-ternary (Apache-2.0)
Container format and engine       github.com/Neroued/ninfer (Apache-2.0)

"Created using Bonsai by Prism ML."

See NOTICE, which also discloses one non-Apache link upstream in the fine-tune family. Apache-2.0 β€” see LICENSE.

"Qwen" is a trademark of Alibaba Cloud; "Bonsai" and "Prism ML" belong to Prism ML, Inc. This is an unofficial, community-produced derivative and is not endorsed by or affiliated with Alibaba Cloud, Prism ML, ukisai, shensanshu, or the NInfer project.

Intended use and limitations

Research and local inference. Not validated for production, safety-critical, or high-stakes use. All figures above were measured in one hardware/software environment and will differ across GPU, driver, CUDA version and memory bandwidth.

δΈ­ζ–‡θ―΄ζ˜Ž