--- license: apache-2.0 base_model: prism-ml/Ternary-Bonsai-2-27B-mlx-2bit pipeline_tag: image-text-to-text library_name: mlx tags: - mlx - jang - affine-quantization - qwen3.8 - bonsai - bonsai-2 - ternary - packed-trits - hadamard - multimodal - vision - video - reasoning - thinking ---

OsaurusAI

# Bonsai-2-27B-1.75bit-JANG JANG-affine bundle of [prism-ml/Ternary-Bonsai-2-27B-mlx-2bit](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit), Prism ML's ternary Bonsai 2 27B (built on Qwen3.8-27B). This is the **dense-packed** edition: the same exact ternary weights as [Bonsai-2-27B-Ternary-JANG](https://huggingface.co/OsaurusAI/Bonsai-2-27B-Ternary-JANG), stored at Prism's PTQ1_0 density (1.75 bits/weight: 5 trits per byte plus one fp16 scale per 128-weight group, no stored bias) and expanded losslessly into MLX's native 2-bit affine kernels at load time. Output is bit-identical to the ternary bundle; only the download shrinks. Proper affine JANG storage, not JANGTQ, MXTQ, or a codebook sidecar format. [OsaurusAI](https://huggingface.co/OsaurusAI) · [osaurus.ai](https://osaurus.ai) · [JANG source](https://github.com/jjang-ai/jangq) ## Bundle | Property | Value | |---|---| | Architecture | Dense Qwen3.8-27B conditional-generation VLM (64 blocks: 48 GatedDeltaNet + 16 full attention) | | JANG profile | `JANG_AFFINE_TERNARY_PACKED` | | Text matrices | ternary {−s, 0, +s}, 26 bytes of packed trits + fp16 scale per 128-group on disk (1.75 bpw), 2-bit slots in memory, exact | | Weight basis | blockwise Hadamard rotation (block 1024, explicit signs), applied to activations at runtime | | Vision linears | 6-bit affine, group size 128 | | Norms and recurrent-state tensors | float32 passthrough (as in the source) | | Weight shards | 6.08 GiB | | Context | 262,144 tokens | | Modalities | text, image, video | | Audio | not supported | Embeddings, the untied language-model head, full-attention projections, GatedDeltaNet projections, and MLP matrices are ternary. Bonsai is dense: it has no routed experts or router tensors. The ternary codes decode to exactly the source {−scale, 0, +scale} groups. The bundle contains the original tokenizer, tokenizer config, the Qwen3.8 chat template (thinking, `reasoning_effort`, tools), image and video processor configs, the Prism `hadamard.json` sidecar, source license and notice. EOS metadata is normalized to `<|im_end|>` (`248046`). ## Runtime The language model is stored in a Hadamard-rotated basis. **Stock `mlx_lm` / `mlx_vlm` loaders return wrong output silently** because they skip the activation transform. Use Osaurus or a vMLX build with JANG Hadamard and ternary-packed support (`osaurus.json` names the minimum Osaurus version); the loader expands the packed trits (codes and scales unchanged, biases materialized as −scale), applies the activation transform from the bundle's declared contract, and refuses to load if any module or sign vector is missing. Runtime memory and speed equal the ternary bundle. ```bash vmlx serve OsaurusAI/Bonsai-2-27B-1.75bit-JANG --host 127.0.0.1 --port 8000 ``` OpenAI-compatible chat requests support text, `image_url`, and `video_url` content parts, tool definitions, and `chat_template_kwargs` for `enable_thinking` and `reasoning_effort` (`low`, `medium`, `xhigh`; default `xhigh`). Sampling defaults follow the Qwen3.8 card that Prism also recommends: thinking mode `temperature 1.0, top_p 0.95, top_k 20`; instruct mode `temperature 0.7, top_p 0.80, top_k 20, presence_penalty 1.5`. ## Verification Verified on 2026-09-17 through the vMLX Python server on an Apple M5 Max with 128 GB unified memory. | Gate | Result | |---|---| | Logit parity vs the ternary bundle | PASS — all 402 expanded modules bit-identical, logits bit-identical on 3 prompts | | Logit parity vs Prism's reference loader (via the ternary bundle) | PASS — argmax agreement 1.0 at every position on 4 prompts, identical greedy continuations | | Single-turn text, thinking off | PASS — `Paris` | | Thinking on (`reasoning_effort=medium`) | PASS — closed think block, correct `391` | | Multi-turn | PASS — exact `ORCHID-4729` recall and combination | | Long context | PASS — buried fact recalled from a 10,655-token prompt | | Image | PASS — red background with centered blue square; green circle plus exact OCR of overlaid text | | Video | PASS — red frames followed by blue frames | | Tool calling | PASS — `get_weather` call emitted and tool result folded into the final answer | The conversion report is included as `jang_affine_report.json`; authoritative per-tensor storage metadata is in `jang_config.json`. ## Quantization notes - 402 language-model modules (embedding, 64 layers, untied head) are the source ternary codes and scales, packed 5 trits per byte (26 bytes per 128-group) without re-quantization; biases are always −scale and are not stored. There is no full-precision source for these weights, so AWQ, imatrix and GPTQ do not apply. - 83 eligible vision linears use native 6-bit affine storage; `blocks.N.mlp.linear_fc2` (input 4304) and the patch/position embeddings stay float16. - 699 norms, GatedDeltaNet state projections, convolutions, biases and incompatible vision tensors pass through in their source precision. - No `tq_packed`, `tq_norms`, `mxtq_bits`, or `jangtq_runtime.safetensors` artifacts are present. ## License and attribution Apache-2.0. See `LICENSE` and `NOTICE.txt`. This repository is a repacked conversion of the linked Prism ML Bonsai 2 checkpoint; the ternary weights are Prism ML's work.