Instructions to use Schiltmans/Ternary-Bonsai-2-27B-DFlash2-ft5-mlx-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Schiltmans/Ternary-Bonsai-2-27B-DFlash2-ft5-mlx-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download Schiltmans/Ternary-Bonsai-2-27B-DFlash2-ft5-mlx-4bit --local-dir Ternary-Bonsai-2-27B-DFlash2-ft5-mlx-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Ternary-Bonsai-2-27B-DFlash2-ft5-mlx-4bit
The prequantized form of Ternary-Bonsai-2-27B-DFlash2-ft5, a DFlash 2 drafter fine-tuned against PrismML's Ternary-Bonsai-2-27B. Same weights, not a new training run. The bf16 drafter is quantized to 4 bits every time a runtime loads it. This repository holds the result of doing that once, so the download is 2.83 times smaller. Every measurement on the bf16 card applies unchanged: the bits that reach the GPU are the same.
| Size | model.safetensors 1,361,734,159 bytes, against 3,848,817,907 for the bf16 file |
| Format | 4-bit, group 64, affine, recorded in mlx-lm's standard quantization block (also written as quantization_config). A second block, bonsai2_prequantized, records the same selection with this project's stricter validation |
| Left in bf16, on purpose | 13 matrices: the ten attention_conv/mlp_conv kernel projections, the selector's hidden projection, and both selector codebooks. They are listed as false entries in the quantization block. This is mlx-dspark's selection rule, and the one every measurement used |
| Source | the bf16 drafter at revision 2a2c2c1e, weights sha256 63399215…55b3; manifest.json has the full record |
Load it
| loader | status |
|---|---|
| mlx-dspark with bonsai2-drafter's patches | served and measured: BONSAI2_DRAFTER=Schiltmans/Ternary-Bonsai-2-27B-DFlash2-ft5-mlx-4bit BONSAI2_DRAFTER_REVISION=c0db7e148bb2fe9782f85ea06c2b438b2a47c326 scripts/serve-bonsai2.sh (v0.2.0 or later of the quickstart; the revision pins this exact upload). Stock mlx-dspark 0.18.0 ignores the quantization block and cannot load this file without the patch |
mlx-lm's load_model, with a DFlash 2 model class |
loads: all 153 tensors identical to the patched loader, the same 36 modules at 4 bits |
dflash-mlx-bonsai2 (load_draft_bundle(path, draft_quant=None)) |
loads: all 153 tensors identical. Do not pass --draft-quant with this file: that runtime would also pack the 13 bf16 matrices, codebooks included |
Both loader checks ran on this upload's config.json and model.safetensors; their records are
in the bonsai2-drafter repository's
evidence/prequantized-4bit/second-loaders/.
Plain mlx_lm.load(path) is not expected to work, since the config's model_type is the drafter
architecture, and the bf16 file behaves the same way.
Validation
- Tensor equality. The saved artifact, loaded through the patch, equals the bf16 drafter
quantized at load time: every parameter name, shape, dtype and value, the module classes, the
quantization parameters and the DFlash 2 configuration. That holds in-process and in a fresh
process that cannot read the source directory (271 checks). The exporter itself refuses to
publish an artifact that does not reload identically. That check is the repository's
bench/tests/test_dflash_prequantized.py --source … --artifact …; its log is not in the evidence bundle. - Served equivalence. On the frozen prompt suites, the bf16 and 4-bit packagings produced the
same served loop: the same tokens, finish reasons, token counts, rounds and round lengths
(contract
artifact-equivalence/v1). HTTP decode ran ABBA at a throughput ratio of 1.0001, with no newly failing quality fixture and 56/56 tool calls in both arms. Those runs used the samemodel.safetensorsbytes under a different metadata layout. This release changes onlyconfig.json, and the tensor file is byte-identical. The records are in the repository'sevidence/prequantized-4bit/. - One Apple M4 Pro, 48 GB. No speed claim is made for this packaging beyond "the same as bf16".
Limitations
Those of the bf16 drafter apply, including the disclosed non-termination of one quality prompt at an 8-bit KV cache, which is a property of the cache setting and not of the drafter. There is no 4-bit convention for DFlash 2 drafters in common use. This one follows mlx-lm's, keeps the codebooks in bf16, and says so in its config.
License
Apache-2.0, with the bf16 release's LICENSE and NOTICE, which carry Prism ML's notice and its
requested attribution: "Created using Bonsai by Prism ML." This is an independent release, not an
official Prism ML, Qwen, Inco AI or z-lab one.
- Downloads last month
- 530
4-bit
Model tree for Schiltmans/Ternary-Bonsai-2-27B-DFlash2-ft5-mlx-4bit
Base model
Qwen/Qwen3.8-27B