Ternary-Bonsai-2-27B-DFlash2-ft5-mlx-4bit

The prequantized form of Ternary-Bonsai-2-27B-DFlash2-ft5, a DFlash 2 drafter fine-tuned against PrismML's Ternary-Bonsai-2-27B. Same weights, not a new training run. The bf16 drafter is quantized to 4 bits every time a runtime loads it. This repository holds the result of doing that once, so the download is 2.83 times smaller. Every measurement on the bf16 card applies unchanged: the bits that reach the GPU are the same.

Size model.safetensors 1,361,734,159 bytes, against 3,848,817,907 for the bf16 file
Format 4-bit, group 64, affine, recorded in mlx-lm's standard quantization block (also written as quantization_config). A second block, bonsai2_prequantized, records the same selection with this project's stricter validation
Left in bf16, on purpose 13 matrices: the ten attention_conv/mlp_conv kernel projections, the selector's hidden projection, and both selector codebooks. They are listed as false entries in the quantization block. This is mlx-dspark's selection rule, and the one every measurement used
Source the bf16 drafter at revision 2a2c2c1e, weights sha256 63399215…55b3; manifest.json has the full record

Load it

loader status
mlx-dspark with bonsai2-drafter's patches served and measured: BONSAI2_DRAFTER=Schiltmans/Ternary-Bonsai-2-27B-DFlash2-ft5-mlx-4bit BONSAI2_DRAFTER_REVISION=c0db7e148bb2fe9782f85ea06c2b438b2a47c326 scripts/serve-bonsai2.sh (v0.2.0 or later of the quickstart; the revision pins this exact upload). Stock mlx-dspark 0.18.0 ignores the quantization block and cannot load this file without the patch
mlx-lm's load_model, with a DFlash 2 model class loads: all 153 tensors identical to the patched loader, the same 36 modules at 4 bits
dflash-mlx-bonsai2 (load_draft_bundle(path, draft_quant=None)) loads: all 153 tensors identical. Do not pass --draft-quant with this file: that runtime would also pack the 13 bf16 matrices, codebooks included

Both loader checks ran on this upload's config.json and model.safetensors; their records are in the bonsai2-drafter repository's evidence/prequantized-4bit/second-loaders/.

Plain mlx_lm.load(path) is not expected to work, since the config's model_type is the drafter architecture, and the bf16 file behaves the same way.

Validation

  • Tensor equality. The saved artifact, loaded through the patch, equals the bf16 drafter quantized at load time: every parameter name, shape, dtype and value, the module classes, the quantization parameters and the DFlash 2 configuration. That holds in-process and in a fresh process that cannot read the source directory (271 checks). The exporter itself refuses to publish an artifact that does not reload identically. That check is the repository's bench/tests/test_dflash_prequantized.py --source … --artifact …; its log is not in the evidence bundle.
  • Served equivalence. On the frozen prompt suites, the bf16 and 4-bit packagings produced the same served loop: the same tokens, finish reasons, token counts, rounds and round lengths (contract artifact-equivalence/v1). HTTP decode ran ABBA at a throughput ratio of 1.0001, with no newly failing quality fixture and 56/56 tool calls in both arms. Those runs used the same model.safetensors bytes under a different metadata layout. This release changes only config.json, and the tensor file is byte-identical. The records are in the repository's evidence/prequantized-4bit/.
  • One Apple M4 Pro, 48 GB. No speed claim is made for this packaging beyond "the same as bf16".

Limitations

Those of the bf16 drafter apply, including the disclosed non-termination of one quality prompt at an 8-bit KV cache, which is a property of the cache setting and not of the drafter. There is no 4-bit convention for DFlash 2 drafters in common use. This one follows mlx-lm's, keeps the codebooks in bf16, and says so in its config.

License

Apache-2.0, with the bf16 release's LICENSE and NOTICE, which carry Prism ML's notice and its requested attribution: "Created using Bonsai by Prism ML." This is an independent release, not an official Prism ML, Qwen, Inco AI or z-lab one.

Downloads last month
530
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Schiltmans/Ternary-Bonsai-2-27B-DFlash2-ft5-mlx-4bit

Base model

Qwen/Qwen3.8-27B
Quantized
(2)
this model

Collection including Schiltmans/Ternary-Bonsai-2-27B-DFlash2-ft5-mlx-4bit