--- language: - en - zh library_name: mlx license: mit pipeline_tag: text-generation base_model: NaiveAI/Naive-N0.5-Flash base_model_relation: quantized tags: - mlx - jang - jangh - quantized - apple-silicon - moe - code - long-context - reasoning - thinking - agent - tool-use - naive_n05_flash - gptq ---

MLX Studio / vMLX — run JANG models on Apple Silicon

Run JANG models in MLX Studio / vMLX

JANGQ-AI

> ⚠️ **Runtime: vMLX (Python) support for this model family is not in a released build yet.** The bundle was built > and validated on a vMLX development build that adds the `naive_n05_flash` family. Released vMLX builds do not include > this architecture and cannot load the bundle. > Known limitation of the development build: prefix / SSD cache reuse is not working for this model yet (answers are > correct; a repeated long prompt is processed again instead of being restored). # JANGQ-AI/Naive-N0.5-Flash-JANGH2 **Naive-N0.5-Flash for 128 GB Macs.** 575 GiB of bf16 weights in **95.89 GiB**, at the speed of a plain MLX quant of the same size and far closer to the original model. A JANGH bundle of [NaiveAI/Naive-N0.5-Flash](https://huggingface.co/NaiveAI/Naive-N0.5-Flash): a 309B MoE (15.5B active, 256 routed experts, top-8) for coding and agentic work, with a native 1M-token context built from sliding-window attention and sparse attention. Text only. - **Routed experts**: JANGH at 2-4 bits (2.54 on average). Codebook-quantized with per-row scales and a blockwise Hadamard rotation, rounded with GPTQ on the full per-expert input statistics. Bits are placed per layer by measurement. - **Everything else**: affine 8-bit, or kept in the source precision (routers, norms, the sparse-attention indexer). | | Top-1 agreement | KL median | |---|---|---| | Ceiling (model vs. itself) | 71.6% | ~0.06 | | **JANGH2** | **64.4%** | **0.170** | | RTN affine, same size | 48.2% | 1.552 | ## Read this first: agreement has a ceiling for this model Agreement with the bf16 model cannot reach 100% here, for any quantization. Naive-N0.5-Flash chooses 8 of 256 experts per token in every layer, and the 8th and 9th candidates are often almost tied. A change far smaller than any quantization error flips some of those choices, and the flips compound over 47 layers. We measured that on the **unquantized** weights: adding relative noise of 0.001 to a single tensor of one layer changes the top-1 token at 28% of the positions. That run is the **ceiling** in every table below. The distance to the ceiling, not the distance to 100%, is what a quantization costs. ## Fidelity vs the bf16 model 68,370 teacher-forced positions on 28 held-out prompts (coding, cybersecurity, agentic, general, Chinese, science, academic; six of them 4-6.5k tokens long), scored against the bf16 model run layer by layer from disk. Top-128 renormalized KL. JANGH2 and the control are served through the runtime with 2048-token prefill chunks. | | Size | median KL ↓ | mean KL ↓ | p90 / p95 / p99 ↓ | top-1 ↑ | top-5 ↑ | top-10 ↑ | |---|---|---|---|---|---|---|---| | Ceiling: bf16 + noise 0.001 on one tensor | 575 GiB | 0.060 | 0.749 | 2.23 / 4.26 / 9.62 | 71.6% | 85.8% | 89.1% | | **Naive-N0.5-Flash-JANGH2** | **95.89 GiB** | **0.170** | **1.008** | **3.11 / 5.37 / 10.50** | **64.4%** | **81.5%** | **85.6%** | | MLX affine RTN, same size ¹ | 95.83 GiB | 1.552 | 2.976 | 8.04 / 11.06 / 16.78 | 48.2% | 64.1% | 68.6% | ¹ A control we built for this comparison: stock MLX affine quantization of the same source, experts at 2-3 bits chosen by the same kind of measurement, non-experts affine 8-bit, no calibration. It is not published. By domain, median KL and top-1 agreement: | domain | positions | ceiling | **JANGH2** | MLX affine RTN | |---|---|---|---|---| | cybersecurity | 6,653 | 0.024 · 76.4% | **0.053 · 73.0%** | 1.653 · 48.5% | | coding | 8,249 | 0.031 · 73.7% | **0.082 · 68.8%** | 1.086 · 53.0% | | coding, long prompts | 9,390 | 0.036 · 74.5% | **0.085 · 69.9%** | 1.165 · 54.3% | | agentic | 7,411 | 0.096 · 66.9% | **0.238 · 60.4%** | 1.864 · 43.7% | | agentic, long prompts | 10,190 | 0.152 · 63.7% | **0.248 · 59.0%** | 2.205 · 48.7% | | general | 5,311 | 0.074 · 71.5% | **0.242 · 60.6%** | 1.864 · 42.8% | | long documents | 10,990 | 0.059 · 74.0% | **0.222 · 63.4%** | 1.465 · 47.1% | | Chinese | 3,158 | 0.038 · 77.8% | **0.172 · 65.1%** | 1.496 · 43.7% | | academic | 3,123 | 0.072 · 73.1% | **0.213 · 62.3%** | 1.542 · 44.8% | | science | 3,895 | 0.100 · 68.5% | **0.356 · 57.2%** | 1.426 · 46.8% | The calibration set is weighted toward coding, tool use and cybersecurity, and that is where the bundle is closest to the original. General, Chinese and science text lose more. ## Tool-use fidelity vs the bf16 model 96 held-out tool conversations (65,719 positions) on tools that never appear in calibration, with 208 "call a tool or answer?" decision points. The bf16 model starts a tool call at 100 of the 112 points where the transcript has one. | | ceiling | **JANGH2** | MLX affine RTN | |---|---|---|---| | tool-call points: model starts a tool call (bf16: 100) | 97 | **101** | 83 | | tool-call points: decisions flipped call ↔ answer vs bf16 | 9 | **7** | 27 | | tool-call points: same next token as bf16 | 92.0% | **93.8%** | 75.9% | | answer points: same next token as bf16 | 91.7% | **88.5%** | 55.2% | | median P(``) at tool-call points (bf16: 0.859) | 0.850 | **0.889** | 0.817 | | lowest P(``) at a tool-call point | 0.053 | **0.133** | 0.0003 | | median KL, whole conversations | 0.292 | **0.395** | 3.112 | | top-1 agreement, whole conversations | 55.1% | **51.3%** | 22.7% | On tool decisions JANGH2 is indistinguishable from the model's own ceiling. These are synthetic agent transcripts, so absolute KL is high in every column; read the columns against each other. The calibration set contains agent-style conversations from the same generator (different tools and wording), so part of this result is in-distribution. ## Live behavior (served, temperature 0) | suite | **JANGH2** | MLX affine RTN | |---|---|---| | 48-case tool-use eval (required / auto × thinking on / off) | **48 / 48** | 34 / 48 | | 20 behavior probes: arithmetic, follow-ups with reasoning history, tool round trips, chained tool calls, multi-line tool arguments, efforts low/high/max, a 5,173-token needle prompt | **20 / 20** | 11 / 20 | Examples of what the control gets wrong: 17 × 19 = "329", "221 = 3 × 731", and a passphrase copied from a long prompt with an extra digit. The bf16 model itself does not fit in 128 GB, so there is no live bf16 column; these suites are easy enough that they check behavior and separate a working bundle from a broken one, not more. ## Speed (M5 Max 128 GB, served, interleaved A/B/A/B, fresh prompts, median of 3) | | **JANGH2** | MLX affine RTN | |---|---|---| | decode (tok/s) | **40.38 / 41.02** | 40.87 / 40.27 | | prefill, ~5.1k-token prompt (tok/s) | **662 / 669** | 635 / 619 | | time to first token, ~5.1k-token prompt | **7.80 s / 7.68 s** | 8.17 s / 8.25 s | | memory after load / peak while serving | 95.9 / 98.9 GiB | 95.8 / 99.0 GiB | | server ready after process start | 20 s | 20 s | Two interleaved runs per bundle, same runtime build, a fresh server for each run, each run the median of 3 probes on a never-seen prompt. Decode is the same within run-to-run variation; long-prompt prefill is **~6% faster** than the plain MLX quant. ## JANGH JANGH is our expert format (it replaced JANGTQ): a fixed blockwise Hadamard-32 transform with no random signs and a two-constant codebook per bit width, stored bit-for-bit in MLX's own packing layout, so a runtime reuses MLX's kernel structure for decode and prefill. For compatibility with runtimes already built against it, the on-disk identifiers keep their original names: `config.json` → `jangtq` block (`version: 2`), per-module `"mode": "jangtq2"`, tensors `*.tq2_packed` / `*.tq2_scales`, and `jang_config.json` → `"format": "jangtq2"`. They are not loadable by, and must not be routed to, JANGTQ v1 loaders. ## What's in the bundle - **Text only.** The source model has no vision, audio or video components. - **Thinking + agentic**: - Thinking is ON by default. The model opens `` itself; with thinking off the template appends ``. - Reasoning efforts are **`low` / `high` / `max`** (default `max`). There is no `medium`: the template renders any other value as Max. - Reasoning is kept in history by default. - **Tool calls**: Qwen3-Coder-style XML (`` / `` / `value`), declared as `tool_parser: xml_function`; tool results use the `tool` role and render as ``. Hermes-style JSON parsers will not work. - **Self-describing**: - `config.json` carries the JANGH format block (codebook, packing, rotation, method) and a per-module `quantization` map (141 expert projections, 197 affine 8-bit modules). - `jang_config.json` records capabilities, calibration and per-layer expert bits. - Raw evaluation results, including the ceiling and the control, are in `evaluation/`. - **Expert bits**: gate/up 2-bit in 28 layers and 3-bit in 19; down 2-bit in 11 layers, 3-bit in 34, 4-bit in 2. - Every shard is alignment-safe (zero-copy memory mapping). 49 shards, 1,148 tensors. ## Serving contract - Sampling: `temperature=1.0, top_p=0.95` (vendor defaults), no repetition penalty - EOS: `151645` · context: 1M native - Reasoning: `reasoning_effort` chat-template kwarg (`low` / `high` / `max`), default `max`; declared as `reasoning_parser: think_xml` - Output layout: the model writes ``, then a blank line, then the answer or the ``. - Memory: prefill long prompts in chunks of at most 2048 tokens. With the weights loaded, a single 6k-token forward needs more GPU memory than a 128 GB machine has left. - The routers must stay in fp32 and the sparse-attention indexer unquantized, as shipped. ## Build details - Source: `NaiveAI/Naive-N0.5-Flash` @ `0235b3b` (bf16, 575 GiB) - Calibration: 618k tokens rendered with the model's own chat template (coding, tool conversations, cybersecurity, agentic, general, Chinese, science, academic), referenced to the bf16 model; evaluation prompts, tools and every tenth calibration document held out - Experts: JANGH (odd-cubic codebook, fp16 per-row scale, Hadamard-32 rotation), GPTQ on all 47 MoE layers, bit allocation measured per layer - Measured and not applied: AWQ, MXFP8 for the non-expert weights, per-row bias correction Quantized and validated by **Jinho Jang** — eric@jangq.ai