LanguaMan pashak commited on
Commit
e27889a
·
0 Parent(s):

Duplicate from prism-ml/Ternary-Bonsai-27B-gguf

Browse files

Co-authored-by: Pasha <pashak@users.noreply.huggingface.co>

.eval_results/aime_2026.yaml ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ - dataset:
2
+ id: MathArena/aime_2026
3
+ task_id: MathArena/aime_2026
4
+ source:
5
+ name: Model Card
6
+ url: https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf
7
+ value: 87.5
.eval_results/gsm8k.yaml ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ - dataset:
2
+ id: openai/gsm8k
3
+ task_id: gsm8k
4
+ source:
5
+ name: Model Card
6
+ url: https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf
7
+ value: 96.06
.eval_results/mmmu_pro.yaml ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ - dataset:
2
+ id: MMMU/MMMU_Pro
3
+ task_id: mmmu_pro_vision
4
+ source:
5
+ name: Model Card
6
+ url: https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf
7
+ value: 68.96
.gitattributes ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ *.7z filter=lfs diff=lfs merge=lfs -text
2
+ *.arrow filter=lfs diff=lfs merge=lfs -text
3
+ *.bin filter=lfs diff=lfs merge=lfs -text
4
+ *.bz2 filter=lfs diff=lfs merge=lfs -text
5
+ *.ckpt filter=lfs diff=lfs merge=lfs -text
6
+ *.ftz filter=lfs diff=lfs merge=lfs -text
7
+ *.gz filter=lfs diff=lfs merge=lfs -text
8
+ *.h5 filter=lfs diff=lfs merge=lfs -text
9
+ *.joblib filter=lfs diff=lfs merge=lfs -text
10
+ *.lfs.* filter=lfs diff=lfs merge=lfs -text
11
+ *.mlmodel filter=lfs diff=lfs merge=lfs -text
12
+ *.model filter=lfs diff=lfs merge=lfs -text
13
+ *.msgpack filter=lfs diff=lfs merge=lfs -text
14
+ *.npy filter=lfs diff=lfs merge=lfs -text
15
+ *.npz filter=lfs diff=lfs merge=lfs -text
16
+ *.onnx filter=lfs diff=lfs merge=lfs -text
17
+ *.ot filter=lfs diff=lfs merge=lfs -text
18
+ *.parquet filter=lfs diff=lfs merge=lfs -text
19
+ *.pb filter=lfs diff=lfs merge=lfs -text
20
+ *.pickle filter=lfs diff=lfs merge=lfs -text
21
+ *.pkl filter=lfs diff=lfs merge=lfs -text
22
+ *.pt filter=lfs diff=lfs merge=lfs -text
23
+ *.pth filter=lfs diff=lfs merge=lfs -text
24
+ *.rar filter=lfs diff=lfs merge=lfs -text
25
+ *.safetensors filter=lfs diff=lfs merge=lfs -text
26
+ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
27
+ *.tar.* filter=lfs diff=lfs merge=lfs -text
28
+ *.tar filter=lfs diff=lfs merge=lfs -text
29
+ *.tflite filter=lfs diff=lfs merge=lfs -text
30
+ *.tgz filter=lfs diff=lfs merge=lfs -text
31
+ *.wasm filter=lfs diff=lfs merge=lfs -text
32
+ *.xz filter=lfs diff=lfs merge=lfs -text
33
+ *.zip filter=lfs diff=lfs merge=lfs -text
34
+ *.zst filter=lfs diff=lfs merge=lfs -text
35
+ *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ Ternary-Bonsai-27B-Q2_0.gguf filter=lfs diff=lfs merge=lfs -text
37
+ Ternary-Bonsai-27B-mmproj-BF16.gguf filter=lfs diff=lfs merge=lfs -text
38
+ Ternary-Bonsai-27B-PQ2_0.gguf filter=lfs diff=lfs merge=lfs -text
39
+ Ternary-Bonsai-27B-Q2_g64.gguf filter=lfs diff=lfs merge=lfs -text
40
+ Ternary-Bonsai-27B-mmproj-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
41
+ Ternary-Bonsai-27B-dspark-bf16.gguf filter=lfs diff=lfs merge=lfs -text
42
+ Ternary-Bonsai-27B-dspark-Q2_0.gguf filter=lfs diff=lfs merge=lfs -text
43
+ Ternary-Bonsai-27B-F16.gguf filter=lfs diff=lfs merge=lfs -text
44
+ Ternary-Bonsai-27B-dspark-Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
LICENSE.txt ADDED
@@ -0,0 +1,177 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Apache License
2
+ Version 2.0, January 2004
3
+ http://www.apache.org/licenses/
4
+
5
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
6
+
7
+ 1. Definitions.
8
+
9
+ "License" shall mean the terms and conditions for use, reproduction,
10
+ and distribution as defined by Sections 1 through 9 of this document.
11
+
12
+ "Licensor" shall mean the copyright owner or entity authorized by
13
+ the copyright owner that is granting the License.
14
+
15
+ "Legal Entity" shall mean the union of the acting entity and all
16
+ other entities that control, are controlled by, or are under common
17
+ control with that entity. For the purposes of this definition,
18
+ "control" means (i) the power, direct or indirect, to cause the
19
+ direction or management of such entity, whether by contract or
20
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
21
+ outstanding shares, or (iii) beneficial ownership of such entity.
22
+
23
+ "You" (or "Your") shall mean an individual or Legal Entity
24
+ exercising permissions granted by this License.
25
+
26
+ "Source" form shall mean the preferred form for making modifications,
27
+ including but not limited to software source code, documentation
28
+ source, and configuration files.
29
+
30
+ "Object" form shall mean any form resulting from mechanical
31
+ transformation or translation of a Source form, including but
32
+ not limited to compiled object code, generated documentation,
33
+ and conversions to other media types.
34
+
35
+ "Work" shall mean the work of authorship, whether in Source or
36
+ Object form, made available under the License, as indicated by a
37
+ copyright notice that is included in or attached to the work
38
+ (an example is provided in the Appendix below).
39
+
40
+ "Derivative Works" shall mean any work, whether in Source or Object
41
+ form, that is based on (or derived from) the Work and for which the
42
+ editorial revisions, annotations, elaborations, or other modifications
43
+ represent, as a whole, an original work of authorship. For the purposes
44
+ of this License, Derivative Works shall not include works that remain
45
+ separable from, or merely link (or bind by name) to the interfaces of,
46
+ the Work and Derivative Works thereof.
47
+
48
+ "Contribution" shall mean any work of authorship, including
49
+ the original version of the Work and any modifications or additions
50
+ to that Work or Derivative Works thereof, that is intentionally
51
+ submitted to Licensor for inclusion in the Work by the copyright owner
52
+ or by an individual or Legal Entity authorized to submit on behalf of
53
+ the copyright owner. For the purposes of this definition, "submitted"
54
+ means any form of electronic, verbal, or written communication sent
55
+ to the Licensor or its representatives, including but not limited to
56
+ communication on electronic mailing lists, source code control systems,
57
+ and issue tracking systems that are managed by, or on behalf of, the
58
+ Licensor for the purpose of discussing and improving the Work, but
59
+ excluding communication that is conspicuously marked or otherwise
60
+ designated in writing by the copyright owner as "Not a Contribution."
61
+
62
+ "Contributor" shall mean Licensor and any individual or Legal Entity
63
+ on behalf of whom a Contribution has been received by Licensor and
64
+ subsequently incorporated within the Work.
65
+
66
+ 2. Grant of Copyright License. Subject to the terms and conditions of
67
+ this License, each Contributor hereby grants to You a perpetual,
68
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
69
+ copyright license to reproduce, prepare Derivative Works of,
70
+ publicly display, publicly perform, sublicense, and distribute the
71
+ Work and such Derivative Works in Source or Object form.
72
+
73
+ 3. Grant of Patent License. Subject to the terms and conditions of
74
+ this License, each Contributor hereby grants to You a perpetual,
75
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
76
+ (except as stated in this section) patent license to make, have made,
77
+ use, offer to sell, sell, import, and otherwise transfer the Work,
78
+ where such license applies only to those patent claims licensable
79
+ by such Contributor that are necessarily infringed by their
80
+ Contribution(s) alone or by combination of their Contribution(s)
81
+ with the Work to which such Contribution(s) was submitted. If You
82
+ institute patent litigation against any entity (including a
83
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
84
+ or a Contribution incorporated within the Work constitutes direct
85
+ or contributory patent infringement, then any patent licenses
86
+ granted to You under this License for that Work shall terminate
87
+ as of the date such litigation is filed.
88
+
89
+ 4. Redistribution. You may reproduce and distribute copies of the
90
+ Work or Derivative Works thereof in any medium, with or without
91
+ modifications, and in Source or Object form, provided that You
92
+ meet the following conditions:
93
+
94
+ (a) You must give any other recipients of the Work or
95
+ Derivative Works a copy of this License; and
96
+
97
+ (b) You must cause any modified files to carry prominent notices
98
+ stating that You changed the files; and
99
+
100
+ (c) You must retain, in the Source form of any Derivative Works
101
+ that You distribute, all copyright, patent, trademark, and
102
+ attribution notices from the Source form of the Work,
103
+ excluding those notices that do not pertain to any part of
104
+ the Derivative Works; and
105
+
106
+ (d) If the Work includes a "NOTICE" text file as part of its
107
+ distribution, then any Derivative Works that You distribute must
108
+ include a readable copy of the attribution notices contained
109
+ within such NOTICE file, excluding those notices that do not
110
+ pertain to any part of the Derivative Works, in at least one
111
+ of the following places: within a NOTICE text file distributed
112
+ as part of the Derivative Works; within the Source form or
113
+ documentation, if provided along with the Derivative Works; or,
114
+ within a display generated by the Derivative Works, if and
115
+ wherever such third-party notices normally appear. The contents
116
+ of the NOTICE file are for informational purposes only and
117
+ do not modify the License. You may add Your own attribution
118
+ notices within Derivative Works that You distribute, alongside
119
+ or as an addendum to the NOTICE text from the Work, provided
120
+ that such additional attribution notices cannot be construed
121
+ as modifying the License.
122
+
123
+ You may add Your own copyright statement to Your modifications and
124
+ may provide additional or different license terms and conditions
125
+ for use, reproduction, or distribution of Your modifications, or
126
+ for any such Derivative Works as a whole, provided Your use,
127
+ reproduction, and distribution of the Work otherwise complies with
128
+ the conditions stated in this License.
129
+
130
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
131
+ any Contribution intentionally submitted for inclusion in the Work
132
+ by You to the Licensor shall be under the terms and conditions of
133
+ this License, without any additional terms or conditions.
134
+ Notwithstanding the above, nothing herein shall supersede or modify
135
+ the terms of any separate license agreement you may have executed
136
+ with Licensor regarding such Contributions.
137
+
138
+ 6. Trademarks. This License does not grant permission to use the trade
139
+ names, trademarks, service marks, or product names of the Licensor,
140
+ except as required for reasonable and customary use in describing the
141
+ origin of the Work and reproducing the content of the NOTICE file.
142
+
143
+ 7. Disclaimer of Warranty. Unless required by applicable law or
144
+ agreed to in writing, Licensor provides the Work (and each
145
+ Contributor provides its Contributions) on an "AS IS" BASIS,
146
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
147
+ implied, including, without limitation, any warranties or conditions
148
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
149
+ PARTICULAR PURPOSE. You are solely responsible for determining the
150
+ appropriateness of using or redistributing the Work and assume any
151
+ risks associated with Your exercise of permissions under this License.
152
+
153
+ 8. Limitation of Liability. In no event and under no legal theory,
154
+ whether in tort (including negligence), contract, or otherwise,
155
+ unless required by applicable law (such as deliberate and grossly
156
+ negligent acts) or agreed to in writing, shall any Contributor be
157
+ liable to You for damages, including any direct, indirect, special,
158
+ incidental, or consequential damages of any character arising as a
159
+ result of this License or out of the use or inability to use the
160
+ Work (including but not limited to damages for loss of goodwill,
161
+ work stoppage, computer failure or malfunction, or any and all
162
+ other commercial damages or losses), even if such Contributor
163
+ has been advised of the possibility of such damages.
164
+
165
+ 9. Accepting Warranty or Additional Liability. While redistributing
166
+ the Work or Derivative Works thereof, You may choose to offer,
167
+ and charge a fee for, acceptance of support, warranty, indemnity,
168
+ or other liability obligations and/or rights consistent with this
169
+ License. However, in accepting such obligations, You may act only
170
+ on Your own behalf and on Your sole responsibility, not on behalf
171
+ of any other Contributor, and only if You agree to indemnify,
172
+ defend, and hold each Contributor harmless for any liability
173
+ incurred by, or claims asserted against, such Contributor by reason
174
+ of your accepting any such warranty or additional liability.
175
+
176
+ END OF TERMS AND CONDITIONS
177
+
NOTICE.txt ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ This software is copyright 2026-present Prism ML, Inc. It is available under the Apache 2.0 license.
2
+ If you publicly deploy or redistribute this software, we would appreciate attribution such as: "Created using Bonsai by Prism ML."
3
+
4
+ This software is built from Qwen3.6-27B, Copyright 2026 Alibaba Cloud, which is available under the Apache 2.0 License: https://huggingface.co/Qwen/Qwen3.6-27B/blob/main/LICENSE
README.md ADDED
@@ -0,0 +1,321 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: llama.cpp
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - conversational
7
+ - ternary
8
+ - 2-bit
9
+ - gguf
10
+ - llama-cpp
11
+ - cuda
12
+ - metal
13
+ - on-device
14
+ - hybrid-attention
15
+ - prismml
16
+ - bonsai
17
+ base_model:
18
+ - Qwen/Qwen3.6-27B
19
+ ---
20
+
21
+ <p align="center">
22
+ <img src="./assets/bonsai-logo.svg" width="280" alt="Bonsai">
23
+ </p>
24
+
25
+ <p align="center">
26
+ <a href="https://prismml.com"><b>Prism ML Website</b></a> &nbsp;|&nbsp;
27
+ <a href="https://github.com/PrismML-Eng/Bonsai-demo"><b>Whitepaper</b></a> &nbsp;|&nbsp;
28
+ <a href="https://github.com/PrismML-Eng/Bonsai-demo"><b>Demo &amp; Examples</b></a> &nbsp;|&nbsp;
29
+ <a href="https://discord.gg/prismml"><b>Discord</b></a>
30
+ </p>
31
+
32
+ # Ternary Bonsai 27B — GGUF
33
+
34
+ Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU)
35
+
36
+ > **\~9.4x** smaller than FP16 (ideal) | **95%** of FP16 intelligence retained | **\~26 tok/s** on an Apple M5 Pro laptop
37
+
38
+ ## Highlights
39
+
40
+ - **\~7.2 GB** deployed footprint (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU
41
+ - **95% of FP16 intelligence retained**: 80.49 average across 15 thinking-mode benchmarks — a *higher* score than the conventional IQ2_XXS build (72.73) at less than two-thirds of its footprint
42
+ - **Retains thinking, reasoning, and agentic behavior** deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within two points of full precision (93.40), coding at 85.96, agentic tool use at 74.01
43
+ - **End-to-end ternary language weights** across embeddings, attention projections, MLP projections, and LM head, at a *true* 1.71 bits per weight — no high-precision escape hatches behind a low-bit label; the vision tower ships in compact 4-bit HQQ
44
+ - **262K-token context** on-device, kept practical by the Qwen3.6-27B hybrid-attention backbone (\~75% linear attention) and 4-bit KV-cache quantization
45
+ - **GGUF Q2_0_g128** format with custom 2-bit hybrid-attention kernels for llama.cpp (CUDA, Metal) — packed weights are consumed directly, never expanded back to FP16
46
+ - **Ships with a DSpark speculative-decoding drafter layer** trained against the Bonsai 27B target — a lossless **1.34x** decode speedup on the CUDA serving path
47
+ - **MLX companion**: also available as [Ternary-Bonsai-27B-mlx-2bit](https://huggingface.co/prism-ml/Ternary-Bonsai-27B-mlx-2bit) for native Apple Silicon inference
48
+ - **1-bit companion**: the phone-class operating point (\~3.9 GB) that fits an iPhone 17 Pro Max, published in GGUF as [Bonsai-27B-gguf](https://huggingface.co/prism-ml/Bonsai-27B-gguf)
49
+
50
+ ## Resources
51
+
52
+ - **[Whitepaper](https://github.com/PrismML-Eng/Bonsai-demo/blob/main/bonsai-27b-whitepaper.pdf)** — full methodology, benchmarks, and measurement notes
53
+ - **[Demo & examples](https://github.com/PrismML-Eng/Bonsai-demo)** — serving, benchmarking, and integrating Bonsai
54
+ - **Low-bit kernels**: [llama.cpp fork](https://github.com/PrismML-Eng/llama.cpp) (CUDA + Metal) · [MLX fork](https://github.com/PrismML-Eng/mlx) (Apple Silicon) · [mlx-swift fork](https://github.com/PrismML-Eng/mlx-swift) (iOS/macOS)
55
+ - **[Discord](https://discord.gg/prismml)** — join the community for support, discussion, and updates
56
+
57
+ ## Model Overview
58
+
59
+ | Item | Specification |
60
+ | :---------------- | :----------------------------------------------------------------------------------------------- |
61
+ | Base model | Derived from Qwen3.6-27B, a 27B hybrid-attention causal language model (architecture unchanged) |
62
+ | Parameters | \~27.3B ternary language weights (\~24.8B backbone across 64 blocks + \~2.5B embedding/LM head) + \~0.46B vision tower (27 blocks) |
63
+ | Architecture | Hybrid attention (\~75% linear / \~25% full attention), SwiGLU MLP, RoPE, RMSNorm |
64
+ | Context length | 262K tokens (full-context capable on-device, enabled by the predominantly linear-attention backbone) |
65
+ | KV cache | Near-lossless 4-bit KV quantization; the hybrid backbone grows a full-attention cache on only 16 of 64 layers (\~4.3 GB at the full 262K window) |
66
+ | Weight format | GGUF Q2_0_g128: {−1, 0, +1} weights in 2-bit slots with FP16 group-wise scaling |
67
+ | Low-bit coverage | Embeddings, attention projections, MLP projections, LM head |
68
+ | Vision tower | HQQ 4-bit; optional \~0.63 GB mmproj pack (Q8_0 container), loaded only for image input |
69
+ | Deployed size | **\~7.2 GB** (5.9 GB ideal at 1.71 bits/weight; see below) |
70
+ | Acceleration | DSpark speculative-decoding drafter layer provided |
71
+ | Backends | llama.cpp (CUDA, Metal, CPU) |
72
+ | License | Apache 2.0 |
73
+
74
+ ## Weight Representation: Q2_0_g128
75
+
76
+ Each weight takes a value from {−1, 0, +1}, with one shared FP16 scale factor for every group of 128 weights. A ternary value carries log₂3 ≈ 1.585 bits of information, so the effective storage cost is **\~1.71 bits/weight** (ternary code + 16-bit scale amortized over 128 weights) — an idealized \~9.4x reduction vs FP16.
77
+
78
+ Relative to the binary format, the extra zero state gives a more expressive weight alphabet and recovers more of the full-precision model's behavior, which makes ternary the **quality-oriented operating point** of the Bonsai 27B family.
79
+
80
+ ### Memory Requirement
81
+
82
+ | Format | True bits/weight | Ideal size | Deployed size | Reduction (ideal) |
83
+ | :------------------ | ---------------: | ---------: | ------------: | ----------------: |
84
+ | FP16 (baseline) | 16.0 | \~54 GB | — | 1.0x |
85
+ | **GGUF Q2_0_g128** | **1.71** | **5.9 GB** | **\~7.2 GB** | **\~9.4x** |
86
+
87
+ Today's kernels store each ternary value in a 2-bit slot (2.125 bits/weight deployed), so the deployed footprint sits above the representation's information-theoretic minimum until native ternary kernels close the gap. The deployed figure describes the language model alone — the only component that must stay resident for text inference; a negligible tail of normalization and scale parameters remains in higher precision.
88
+
89
+ Unlike conventional low-bit builds — whose advertised labels understate their true average bit-width (a widely-used "2-bit" build of Qwen3.6-27B is really 2.8 bits/weight at 9.4 GB) — the Bonsai representation carries a bit-width that matches its name.
90
+
91
+ ### Shipped Components
92
+
93
+ Two optional components ship alongside the language model (on-disk sizes):
94
+
95
+ | Component | Pack | Size | Residency |
96
+ | :------------- | :--------------------------------- | -------: | :--------------------------------- |
97
+ | Language model | 2-bit g128 slots (Q2_0) | 7.17 GB | resident |
98
+ | DSpark drafter | Q4_1 (default) | 1.95 GB | optional — speculative decoding |
99
+ | DSpark drafter | bf16 (reference) | 7.29 GB | optional |
100
+ | Vision tower | mmproj HQQ 4-bit (Q8_0 container) | 0.63 GB | optional — multimodal input only |
101
+ | Vision tower | mmproj BF16 (reference) | 0.93 GB | optional |
102
+
103
+ The vision tower is usually offloaded: it sits outside the accelerator's resident budget and is loaded only when an image actually arrives, so text-only serving never pays for it. A group-64 ternary pack (7.59 GB) is also published, matching the 64-value-group Q2_0 packing in llama.cpp — the same native g128 representation with each scale repeated per 64-value block.
104
+
105
+ ### Peak Memory at Context
106
+
107
+ What a device must actually accommodate is *peak* memory — weights plus KV cache plus activations and runtime buffers (\~1.3 GB across backends). Measured, language model only, no KV-cache compression (sizes in decimal GB; the Q4_K_XL row is derived from its weight footprint plus the same measured cache-and-overhead build-up, all other rows directly measured):
108
+
109
+ | Build | Weights | 4K ctx | 10K ctx | 100K ctx |
110
+ | :----------------------------------- | ------: | -----: | ------: | -------: |
111
+ | **Ternary Bonsai (llama.cpp Q2_0)** | 7.15 | 8.4 | 8.7 | 14.7 |
112
+ | Qwen3.6-27B "4-bit" (Q4_K_XL) | 17.6 | 19.2 | 19.6 | 25.6 |
113
+ | 27B 16-bit (GGUF bf16) | 51.25 | 52.6 | 53.3 | 59.3 |
114
+
115
+ The ternary build holds a **100K-token context at 14.7 GB without any KV-cache compression** — a budget that fits mainstream laptops outright; the conventional Q4_K_XL build needs \~25.6 GB before the first long document is loaded. These peaks are the conservative case, with the cache left at FP16. Enabling the 4-bit KV cache shrinks the context-dependent term \~4x: the 100K peak drops to \~10.1 GB, and the full 262K window fits in \~12.8 GB peak.
116
+
117
+ ## Best Practices
118
+
119
+ ### Generation Parameters
120
+
121
+ | Parameter | Suggested |
122
+ | :---------- | :-------- |
123
+ | Temperature | 0.7 |
124
+ | Top-p | 0.95 |
125
+ | Top-k | 20 |
126
+
127
+ These are the settings used for all reported benchmark results (thinking mode).
128
+
129
+ ### System Prompt
130
+
131
+ You can use a simple system prompt such as:
132
+
133
+ ```
134
+ You are a helpful assistant
135
+ ```
136
+
137
+ ## Quickstart
138
+
139
+ ### llama.cpp (CUDA)
140
+
141
+ ```bash
142
+ # Clone the PrismML fork of llama.cpp (includes the Q2_0_g128 hybrid-attention kernels)
143
+ git clone https://github.com/PrismML-Eng/llama.cpp
144
+ cd llama.cpp
145
+
146
+ # Build with CUDA support
147
+ cmake -B build -DGGML_CUDA=ON && cmake --build build -j
148
+
149
+ # Download the 2-bit GGUF weights
150
+ hf download prism-ml/Ternary-Bonsai-27B-gguf Ternary-Bonsai-27B-Q2_0.gguf --local-dir .
151
+
152
+ # Run inference
153
+ ./build/bin/llama-cli \
154
+ -m Ternary-Bonsai-27B-Q2_0.gguf \
155
+ -p "Explain quantum computing in simple terms." \
156
+ -n 256 \
157
+ --temp 0.7 --top-p 0.95 --top-k 20 \
158
+ -ngl 99
159
+ ```
160
+
161
+ ### llama.cpp (Metal / macOS)
162
+
163
+ ```bash
164
+ # Build with Metal support (default on macOS)
165
+ cmake -B build && cmake --build build -j
166
+
167
+ # Run inference
168
+ ./build/bin/llama-cli \
169
+ -m Ternary-Bonsai-27B-Q2_0.gguf \
170
+ -p "Explain quantum computing in simple terms." \
171
+ -n 256 \
172
+ --temp 0.7 --top-p 0.95 --top-k 20 \
173
+ -ngl 99
174
+ ```
175
+
176
+ ### llama.cpp Server
177
+
178
+ ```bash
179
+ ./build/bin/llama-server \
180
+ -m Ternary-Bonsai-27B-Q2_0.gguf \
181
+ --host 0.0.0.0 --port 8080 -ngl 99
182
+ ```
183
+
184
+ Open the web UI at [http://127.0.0.1:8080](http://127.0.0.1:8080), or see our [llama.cpp fork](https://github.com/PrismML-Eng/llama.cpp) for more examples.
185
+
186
+ > **Deploying to a phone?** The ternary build (\~7.2 GB) exceeds the \~6 GB per-app iOS memory budget and is laptop/GPU-only. Use the 1-bit companion (\~3.9 GB), which fits an iPhone 17 Pro Max via [MLX Swift](https://huggingface.co/prism-ml/Bonsai-27B-mlx-1bit).
187
+
188
+ ## Cross-Platform Throughput
189
+
190
+ `tg128` is token-generation throughput over 128 generated tokens (the memory-bandwidth-bound, interactive phase); `pp512` is prompt-processing throughput over 512 input tokens (the compute-bound phase). Both in tokens/s, measured with `llama-bench` on this GGUF pack (custom low-bit kernels).
191
+
192
+ | Platform | Footprint | TG128 (tok/s) | PP512 (tok/s) |
193
+ | :--------------------------- | --------: | ------------: | ------------: |
194
+ | Laptop (Apple M5 Max, Metal) | 7.2 GB | 44.0 | 830 |
195
+ | Laptop (Apple M5 Pro, Metal) | 7.2 GB | 26.2 | 393 |
196
+ | Laptop (Apple M4 Pro, Metal) | 7.2 GB | 18.0 | 125 |
197
+ | Single GPU (H100, CUDA) | 7.2 GB | 98.0 | 2596 |
198
+
199
+ On the laptop, the FP16 baseline (\~54 GB) and even conventional "4-bit" builds (17.6 GB) do not fit at all — the meaningful statement is not a speedup ratio but that a 27B model runs interactively on an everyday laptop. The measured decode streams \~186 GB/s of weights on the M5 Pro, confirming the memory-bandwidth-dominated profile that the low-bit representation is built to exploit. The H100 row is the exception that proves the rule: at batch size 1 a datacenter GPU is limited by kernel-launch and synchronization latency rather than weight bandwidth, so the ternary and binary variants converge there (98 vs 104.8 tok/s) despite their \~1.9x difference in bytes per step.
200
+
201
+ ## Speculative Decoding: DSpark
202
+
203
+ Ternary Bonsai 27B ships with a **DSpark** drafter layer trained against the low-bit target — a semi-autoregressive drafter with confidence-scheduled verification. Speculative decoding is lossless: verification preserves the target distribution exactly, so accepted tokens are indistinguishable from ordinary generation.
204
+
205
+ The drafter is a compact **six-layer block-parallel transformer** conditioned on hidden states tapped from five evenly spaced layers of the target; its drafter-unique weights add roughly **0.5 GB at serving precision** (embeddings and output head are shared with the resident target). It follows the DSpark recipe with a diffusion-flavored block-denoising objective, survival-probability-weighted distillation, per-source-normalized hidden-state taps, and a draft block size chosen from a measured verify-cost model of the serving stack. The drafter ships 4-bit quantized — the \~1.95 GB Q4_1 pack is the default; it drafts faster than the bf16 reference at essentially unchanged draft quality, and because verification preserves the target distribution exactly, drafter precision affects only speed, never output quality.
206
+
207
+ On the CUDA serving path the drafter is a measured net win — an accepted length of τ ≈ 3.7 at draft depth k = 4 turns into a **1.34x** end-to-end decode speedup on H100 (98 → 131.8 tok/s). On Apple Silicon the batch-1 verification pass does not yet amortize, so the drafter layer is not enabled by default on-device.
208
+
209
+ ## Benchmarks
210
+
211
+ Evaluated with EvalScope + vLLM on NVIDIA H100 under identical infrastructure, decoding, and scoring, in **thinking mode** — where the model's full reasoning is exercised and the sub-4-bit collapse of conventional methods is most visible. 15 benchmarks across six skill categories. For cross-family context the table also includes Gemma-4-31B, a model of the same capability tier, with its conventional low-bit builds — the collapse below 4 bits is a property of the methods, not of one base model. Bit-widths are true averages; "vs FP16" is relative to the Qwen3.6-27B FP16 reference.
212
+
213
+ | Variant | True bpw | Footprint | Thinking avg | vs FP16 |
214
+ | :-------------------------------------------------------------------------- | -------: | ---------: | -----------: | ---------: |
215
+ | Qwen3.6-27B FP16 | 16.0 | 54 GB | 85.07 | 100% |
216
+ | Qwen3.6-27B Q4_K_XL ("4-bit") | 5.2 | 17.6 GB | 84.99 | 99.9% |
217
+ | Qwen3.6-27B IQ2_XXS ("2-bit") | 2.8 | 9.4 GB | 72.73 | 85.5% |
218
+ | Gemma-4-31B FP16 | 16.0 | 61.5 GB | 84.58 | 99.4% |
219
+ | Gemma-4-31B QAT ("4-bit") | 6.0 | 23.3 GB | 83.41 | 98.0% |
220
+ | Gemma-4-31B Q2_K_XL ("2-bit") | 3.0 | 11.8 GB | 73.31 | 86.2% |
221
+ | **Ternary Bonsai 27B** | **1.71** | **5.9 GB** | **80.49** | **94.6%** |
222
+ | 1-bit Bonsai 27B | 1.125 | 3.9 GB | 76.11 | 89.5% |
223
+
224
+ At 5.9 GB, Ternary Bonsai 27B outscores both sub-4-bit conventional builds by more than seven points at one-half to two-thirds of their size.
225
+
226
+ The aggregate gap also understates *how* the conventional builds fail: their degradation is selective, concentrated on the benchmarks that demand sustained chains of reasoning. IQ2_XXS falls to 57.5 on AIME26 and 56.4 on LiveCodeBench while still scoring 88.93 on MMLU-Redux — which is why casual testing misses the collapse. Ternary Bonsai holds exactly these benchmarks, keeping AIME at 87.5–90.8 and LiveCodeBench at 82.8.
227
+
228
+ ### By Skill Category
229
+
230
+ | Category | Benchmarks | FP16 | Ternary 27B |
231
+ | :---------------------- | :---------------------------------- | ----: | ----------: |
232
+ | Knowledge & reasoning | MMLU-Redux, MuSR | 83.15 | 76.96 |
233
+ | Math | GSM8K, MATH-500, AIME25, AIME26 | 95.33 | 93.40 |
234
+ | Coding | HumanEval+, MBPP+, LiveCodeBench | 88.74 | 85.96 |
235
+ | Instruction following | IFEval, IFBench | 78.47 | 71.77 |
236
+ | Agentic / tool calling | BFCL v3, τ²-Bench | 80.00 | 74.01 |
237
+ | Vision | MMMU-Pro, OCR Bench v2 | 72.61 | 65.19 |
238
+ | **Overall (15)** | | **85.07** | **80.49** |
239
+
240
+ The reasoning backbone comes through intact: math stays within two points of full precision (93.40), coding at 85.96, and the ternary model spends its extra footprint to hold the most demanding categories — agentic tool use at 74.01 and vision at 65.19 — the behaviors that conventional sub-4-bit representations lose first.
241
+
242
+ ### Full Per-Benchmark Results
243
+
244
+ <details>
245
+ <summary>Expand full per-benchmark results (thinking mode)</summary>
246
+
247
+ | Benchmark | FP16 | Ternary 27B |
248
+ | :--------------------- | ----: | ----------: |
249
+ | MMLU-Redux | 93.42 | 88.05 |
250
+ | MuSR | 72.88 | 65.87 |
251
+ | GSM8K | 95.30 | 96.06 |
252
+ | MATH-500 | 99.40 | 99.20 |
253
+ | AIME25 | 93.29 | 90.84 |
254
+ | AIME26 | 93.33 | 87.50 |
255
+ | HumanEval+ | 95.12 | 93.90 |
256
+ | MBPP+ | 83.33 | 81.22 |
257
+ | LiveCodeBench | 87.77 | 82.75 |
258
+ | IFEval | 88.91 | 85.03 |
259
+ | IFBench (prompt-loose) | 68.03 | 58.50 |
260
+ | BFCL v3 | 77.10 | 74.41 |
261
+ | τ²-Bench | 82.90 | 73.61 |
262
+ | MMMU-Pro | 79.94 | 68.96 |
263
+ | OCR Bench v2 | 65.28 | 61.42 |
264
+ | **Average (15)** | **85.07** | **80.49** |
265
+
266
+ </details>
267
+
268
+ ## Intelligence Density
269
+
270
+ Intelligence density captures the ratio of a model's capability to its deployed size:
271
+
272
+ ```
273
+ D = -log2(1 - score/100) / size_GB
274
+ ```
275
+
276
+ | Variant | Size (GB) | Benchmark avg | Intelligence Density (1/GB) |
277
+ | :-------------------------------------------------------------------------- | --------: | -----------: | --------------------------: |
278
+ | 1-bit Bonsai 27B | 3.9 | 76.11 | 0.530 |
279
+ | **Ternary Bonsai 27B** | **5.9** | 80.49 | **0.400** |
280
+ | Qwen3.6-27B IQ2_XXS | 9.4 | 72.73 | 0.199 |
281
+ | Gemma-4-31B Q2_K_XL | 11.8 | 73.31 | 0.162 |
282
+ | Qwen3.6-27B Q4_K_XL | 17.6 | 84.99 | 0.155 |
283
+ | Gemma-4-31B QAT | 23.3 | 83.41 | 0.111 |
284
+ | Qwen3.6-27B FP16 | 54 | 85.07 | 0.051 |
285
+ | Gemma-4-31B FP16 | 61.5 | 84.58 | 0.044 |
286
+
287
+ Ternary Bonsai 27B delivers **2x** the density of the densest conventional build (IQ2_XXS at 0.199) and nearly **8x** FP16 — no conventional build of Qwen3.6-27B or Gemma-4-31B exceeds 0.2. Each stored gigabyte is translated into far more usable intelligence.
288
+
289
+ ## Use Cases
290
+
291
+ - **Laptop-local 27B agents**: full 27B reasoning and tool use on a standard laptop at \~26 tok/s, with the 262K context available for long-document analysis, full-repository code work, and other tasks that depend on holding a large working set in context
292
+ - **Privacy-sensitive and offline settings**: on-device execution keeps prompts and data on the device by construction, and works with intermittent or no connectivity
293
+ - **Single-GPU and commodity-GPU serving**: 27B-class quality from a single consumer or entry-level datacenter GPU, with headroom for larger batches, longer contexts, or co-resident models — combined with the KV-cache quantization, high-throughput serving and long-context document analysis become practical on a single 24 GB GPU
294
+ - **Quality-first low-bit deployment**: when the deployment target has laptop-class memory or better, ternary is the operating point that retains the most of the full-precision model's behavior
295
+
296
+ ## Limitations
297
+
298
+ - **The quality–footprint trade-off**: the ternary model retains 94.6% of the full-precision average, and the gap is modest and predictable — the reasoning core (math, coding) stays within a few points of baseline, with the difference concentrated in the most demanding categories
299
+ - **Does not fit a phone**: at \~7.2 GB the ternary build exceeds the \~6 GB per-app iOS memory budget; use the 1-bit companion via [MLX Swift](https://huggingface.co/prism-ml/Bonsai-27B-mlx-1bit) for phone deployment
300
+ - **Served in 2-bit slots today**: the deployed footprint (\~7.2 GB) sits above the representation's \~5.9 GB native target; native ternary kernels are an active engineering target and would return the remaining bandwidth and footprint advantage directly as latency and energy improvements
301
+ - **Agentic coding** (long-horizon, multi-file, run-test-and-repair workflows) is not yet a strong target of this release; a Bonsai 27B variant tuned for agentic coding is next on the roadmap
302
+ - **KV compression headroom**: this release standardizes on a 4-bit KV cache; early results show the key cache can be pushed toward the sub-2-bit regime — a path to still longer contexts within a fixed device-memory budget
303
+
304
+ ## Citation
305
+
306
+ If you use Ternary Bonsai 27B, please cite:
307
+
308
+ ```bibtex
309
+ @techreport{bonsai27b,
310
+ title = {Bonsai 27B: Full 27B-Class Reasoning in Binary and Ternary
311
+ Transformer Weights --- on Laptops and Phones},
312
+ author = {Prism ML},
313
+ year = {2026},
314
+ month = {July},
315
+ url = {https://prismml.com}
316
+ }
317
+ ```
318
+
319
+ ## Contact
320
+
321
+ For questions, feedback, or collaboration inquiries: **contact@prismml.com**
Ternary-Bonsai-27B-F16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f659ca3dd7e28ada5d8b5f3637862d0d51ef433bde032ec4c8990ed27c91a385
3
+ size 53808280640
Ternary-Bonsai-27B-PQ2_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e4781999f1997ef97ce0c58d05750835acc999d18d83ee6489ba7ac7b14cb5f6
3
+ size 7165121600
Ternary-Bonsai-27B-Q2_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:868c11714cf8fe47f5ec9eeb2be0ab1a337112886f92ee0ede6b855c4fa31757
3
+ size 7165121600
Ternary-Bonsai-27B-Q2_g64.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:59a45d1ecef702b14531b06d22949f33b25c1897da31a8c0b298e01e4d9138eb
3
+ size 7585330240
Ternary-Bonsai-27B-dspark-Q4_1.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c4810091d244eddc61a0cc4966e584b0959f141e3c66c0d371a6652d9f647da9
3
+ size 1946393568
Ternary-Bonsai-27B-dspark-bf16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d5ce05b0e7e23804279fb0b451e330e71c06d977657802a0e22d71433f30dbad
3
+ size 7291885792
Ternary-Bonsai-27B-mmproj-BF16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:acaf5b55d24ebd38c71fa220dc58c9a36776ec543b17728a4b322fc8d92f1de4
3
+ size 931145760
Ternary-Bonsai-27B-mmproj-Q8_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:eb561d41a7bbeb0fcf04883c8af11078ef6cae0a66862a0b68443cfca495269d
3
+ size 629246880
assets/bonsai-logo.svg ADDED