MLX Studio / vMLX — run JANG models on Apple Silicon

Run JANG models in MLX Studio / vMLX

JANGQ-AI

⚠️ Runtime: vMLX (Python) support for this model family is not in a released build yet. The bundle was built and validated on a vMLX development build that adds the naive_n05_flash family. Released vMLX builds do not include this architecture and cannot load the bundle. Known limitation of the development build: prefix / SSD cache reuse is not working for this model yet (answers are correct; a repeated long prompt is processed again instead of being restored).

JANGQ-AI/Naive-N0.5-Flash-JANGH2

Naive-N0.5-Flash for 128 GB Macs. 575 GiB of bf16 weights in 95.89 GiB, at the speed of a plain MLX quant of the same size and far closer to the original model.

A JANGH bundle of NaiveAI/Naive-N0.5-Flash: a 309B MoE (15.5B active, 256 routed experts, top-8) for coding and agentic work, with a native 1M-token context built from sliding-window attention and sparse attention. Text only.

  • Routed experts: JANGH at 2-4 bits (2.54 on average). Codebook-quantized with per-row scales and a blockwise Hadamard rotation, rounded with GPTQ on the full per-expert input statistics. Bits are placed per layer by measurement.
  • Everything else: affine 8-bit, or kept in the source precision (routers, norms, the sparse-attention indexer).
Top-1 agreement KL median
Ceiling (model vs. itself) 71.6% ~0.06
JANGH2 64.4% 0.170
RTN affine, same size 48.2% 1.552

Read this first: agreement has a ceiling for this model

Agreement with the bf16 model cannot reach 100% here, for any quantization. Naive-N0.5-Flash chooses 8 of 256 experts per token in every layer, and the 8th and 9th candidates are often almost tied. A change far smaller than any quantization error flips some of those choices, and the flips compound over 47 layers.

We measured that on the unquantized weights: adding relative noise of 0.001 to a single tensor of one layer changes the top-1 token at 28% of the positions. That run is the ceiling in every table below. The distance to the ceiling, not the distance to 100%, is what a quantization costs.

Fidelity vs the bf16 model

68,370 teacher-forced positions on 28 held-out prompts (coding, cybersecurity, agentic, general, Chinese, science, academic; six of them 4-6.5k tokens long), scored against the bf16 model run layer by layer from disk. Top-128 renormalized KL. JANGH2 and the control are served through the runtime with 2048-token prefill chunks.

Size median KL ↓ mean KL ↓ p90 / p95 / p99 ↓ top-1 ↑ top-5 ↑ top-10 ↑
Ceiling: bf16 + noise 0.001 on one tensor 575 GiB 0.060 0.749 2.23 / 4.26 / 9.62 71.6% 85.8% 89.1%
Naive-N0.5-Flash-JANGH2 95.89 GiB 0.170 1.008 3.11 / 5.37 / 10.50 64.4% 81.5% 85.6%
MLX affine RTN, same size ¹ 95.83 GiB 1.552 2.976 8.04 / 11.06 / 16.78 48.2% 64.1% 68.6%

¹ A control we built for this comparison: stock MLX affine quantization of the same source, experts at 2-3 bits chosen by the same kind of measurement, non-experts affine 8-bit, no calibration. It is not published.

By domain, median KL and top-1 agreement:

domain positions ceiling JANGH2 MLX affine RTN
cybersecurity 6,653 0.024 · 76.4% 0.053 · 73.0% 1.653 · 48.5%
coding 8,249 0.031 · 73.7% 0.082 · 68.8% 1.086 · 53.0%
coding, long prompts 9,390 0.036 · 74.5% 0.085 · 69.9% 1.165 · 54.3%
agentic 7,411 0.096 · 66.9% 0.238 · 60.4% 1.864 · 43.7%
agentic, long prompts 10,190 0.152 · 63.7% 0.248 · 59.0% 2.205 · 48.7%
general 5,311 0.074 · 71.5% 0.242 · 60.6% 1.864 · 42.8%
long documents 10,990 0.059 · 74.0% 0.222 · 63.4% 1.465 · 47.1%
Chinese 3,158 0.038 · 77.8% 0.172 · 65.1% 1.496 · 43.7%
academic 3,123 0.072 · 73.1% 0.213 · 62.3% 1.542 · 44.8%
science 3,895 0.100 · 68.5% 0.356 · 57.2% 1.426 · 46.8%

The calibration set is weighted toward coding, tool use and cybersecurity, and that is where the bundle is closest to the original. General, Chinese and science text lose more.

Tool-use fidelity vs the bf16 model

96 held-out tool conversations (65,719 positions) on tools that never appear in calibration, with 208 "call a tool or answer?" decision points. The bf16 model starts a tool call at 100 of the 112 points where the transcript has one.

ceiling JANGH2 MLX affine RTN
tool-call points: model starts a tool call (bf16: 100) 97 101 83
tool-call points: decisions flipped call ↔ answer vs bf16 9 7 27
tool-call points: same next token as bf16 92.0% 93.8% 75.9%
answer points: same next token as bf16 91.7% 88.5% 55.2%
median P(<tool_call>) at tool-call points (bf16: 0.859) 0.850 0.889 0.817
lowest P(<tool_call>) at a tool-call point 0.053 0.133 0.0003
median KL, whole conversations 0.292 0.395 3.112
top-1 agreement, whole conversations 55.1% 51.3% 22.7%

On tool decisions JANGH2 is indistinguishable from the model's own ceiling. These are synthetic agent transcripts, so absolute KL is high in every column; read the columns against each other. The calibration set contains agent-style conversations from the same generator (different tools and wording), so part of this result is in-distribution.

Live behavior (served, temperature 0)

suite JANGH2 MLX affine RTN
48-case tool-use eval (required / auto × thinking on / off) 48 / 48 34 / 48
20 behavior probes: arithmetic, follow-ups with reasoning history, tool round trips, chained tool calls, multi-line tool arguments, efforts low/high/max, a 5,173-token needle prompt 20 / 20 11 / 20

Examples of what the control gets wrong: 17 × 19 = "329", "221 = 3 × 731", and a passphrase copied from a long prompt with an extra digit. The bf16 model itself does not fit in 128 GB, so there is no live bf16 column; these suites are easy enough that they check behavior and separate a working bundle from a broken one, not more.

Speed (M5 Max 128 GB, served, interleaved A/B/A/B, fresh prompts, median of 3)

JANGH2 MLX affine RTN
decode (tok/s) 40.38 / 41.02 40.87 / 40.27
prefill, ~5.1k-token prompt (tok/s) 662 / 669 635 / 619
time to first token, ~5.1k-token prompt 7.80 s / 7.68 s 8.17 s / 8.25 s
memory after load / peak while serving 95.9 / 98.9 GiB 95.8 / 99.0 GiB
server ready after process start 20 s 20 s

Two interleaved runs per bundle, same runtime build, a fresh server for each run, each run the median of 3 probes on a never-seen prompt. Decode is the same within run-to-run variation; long-prompt prefill is ~6% faster than the plain MLX quant.

JANGH

JANGH is our expert format (it replaced JANGTQ): a fixed blockwise Hadamard-32 transform with no random signs and a two-constant codebook per bit width, stored bit-for-bit in MLX's own packing layout, so a runtime reuses MLX's kernel structure for decode and prefill.

For compatibility with runtimes already built against it, the on-disk identifiers keep their original names: config.json → jangtq block (version: 2), per-module "mode": "jangtq2", tensors *.tq2_packed / *.tq2_scales, and jang_config.json → "format": "jangtq2". They are not loadable by, and must not be routed to, JANGTQ v1 loaders.

What's in the bundle

  • Text only. The source model has no vision, audio or video components.
  • Thinking + agentic:
    • Thinking is ON by default. The model opens <think> itself; with thinking off the template appends <think></think>.
    • Reasoning efforts are low / high / max (default max). There is no medium: the template renders any other value as Max.
    • Reasoning is kept in history by default.
  • Tool calls: Qwen3-Coder-style XML (<tool_call> / <function=name> / <parameter=key>value</parameter>), declared as tool_parser: xml_function; tool results use the tool role and render as <tool_response>. Hermes-style JSON parsers will not work.
  • Self-describing:
    • config.json carries the JANGH format block (codebook, packing, rotation, method) and a per-module quantization map (141 expert projections, 197 affine 8-bit modules).
    • jang_config.json records capabilities, calibration and per-layer expert bits.
    • Raw evaluation results, including the ceiling and the control, are in evaluation/.
  • Expert bits: gate/up 2-bit in 28 layers and 3-bit in 19; down 2-bit in 11 layers, 3-bit in 34, 4-bit in 2.
  • Every shard is alignment-safe (zero-copy memory mapping). 49 shards, 1,148 tensors.

Serving contract

  • Sampling: temperature=1.0, top_p=0.95 (vendor defaults), no repetition penalty
  • EOS: 151645 · context: 1M native
  • Reasoning: reasoning_effort chat-template kwarg (low / high / max), default max; declared as reasoning_parser: think_xml
  • Output layout: the model writes </think>, then a blank line, then the answer or the <tool_call>.
  • Memory: prefill long prompts in chunks of at most 2048 tokens. With the weights loaded, a single 6k-token forward needs more GPU memory than a 128 GB machine has left.
  • The routers must stay in fp32 and the sparse-attention indexer unquantized, as shipped.

Build details

  • Source: NaiveAI/Naive-N0.5-Flash @ 0235b3b (bf16, 575 GiB)
  • Calibration: 618k tokens rendered with the model's own chat template (coding, tool conversations, cybersecurity, agentic, general, Chinese, science, academic), referenced to the bf16 model; evaluation prompts, tools and every tenth calibration document held out
  • Experts: JANGH (odd-cubic codebook, fp16 per-row scale, Hadamard-32 rotation), GPTQ on all 47 MoE layers, bit allocation measured per layer
  • Measured and not applied: AWQ, MXFP8 for the non-expert weights, per-row bias correction

Quantized and validated by Jinho Jang — eric@jangq.ai

Downloads last month
244
Safetensors
Model size
30B params
Tensor type
U32
·
F32
·
BF16
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JANGQ-AI/Naive-N0.5-Flash-JANGH2

Quantized
(4)
this model