MLX Studio / vMLX — run JANG models on Apple Silicon

Run JANG models in MLX Studio / vMLX

JANGQ-AI

✅ Runtime: supported in vMLX (Python). Use vMLX 1.6.68 or newer (native JANGH loading; the vMLX 1.6.71 release was validated with GLM JANGH on native reasoning efforts and chat tool/media continuations). Older vMLX builds refuse the bundle at load time rather than producing wrong output. Known limitation (vMLX 1.6.71 notes): video perception is imperfect — the model can describe overlay or transition text that is not in the clip.

JANGQ-AI/GLM-5.3-Flash-JANGH2

GLM-5.3-Flash for 128 GB Macs. Same size as our previous affine release, with 2.9x lower median KL, +5.3 points top-1, and faster decode.

A JANGH bundle of zai-org/GLM-5.3-Flash: a 300B-class MoE (288 routed experts, top-8 + shared expert) with KDA linear attention, sparse attention, and vision + video towers.

  • Routed experts: JANGH at 2-3 bits. Codebook-quantized with per-row scales and a blockwise Hadamard rotation, calibrated with GPTQ + a per-expert importance matrix. Bits are placed per layer by measurement.
  • Everything else: 8-bit (MXFP8 where the source weights sit on an FP8-MX grid, affine 8-bit elsewhere) or full precision.
  • Vision tower: kept in bf16.

Revision 2: long reasoning now ends on its own

The first upload could not end long reasoning. After a few thousand reasoning tokens the probability of </think> was hundreds to thousands of times too low, so the model kept thinking, and the text eventually degraded into repetition. With the default effort (max) that range is reached on ordinary tasks.

Cause: error-minimizing row scales shrink every low-bit matrix a little (12% at 2 bits). Three matrices per expert and 42 MoE layers deep, that weakens exactly the strong, late decisions, and "stop thinking" is one of them. Revision 2 changes only the row scales of the routed experts (126 small tensors). The quantized codes, the size, the format and the speed are unchanged, and no runtime change is needed.

where the reference model ends its reasoning bf16 revision 2 revision 1 affine JANG (previous)
P(</think>), held-out design task, after 8,025 reasoning tokens 0.82 0.75 0.03 0.02
P(</think>), coding task, after 3,288 reasoning tokens 0.99 0.94 0.19 -
long design prompt, served, vendor sampling: reasoning ended by the model revision 2 revision 1
effort low 3 / 3 0 / 5
effort high 2 / 2 -
effort max (40k-token budget; the reference model needs 17-23k reasoning tokens here) 1 / 1 0 / 2

Your copy is revision 2 if jang_config.json has a scale_correction entry.

This replaces JANGQ-AI/GLM-5.3-Flash-JANG and -JANG-MTP (and was briefly published as GLM-5.3-Flash-JANGTQ2).

JANGH vs JANGTQ

JANGH replaces our earlier JANGTQ (TurboQuant) expert format. The TurboQuant rotation is no longer used: no seeded random-sign rotation and no per-dimension Beta-distribution codebook, so runtimes do not need to reproduce any RNG or dimension-specific tables. JANGH uses a fixed blockwise Hadamard-32 transform (no random signs) and a two-constant codebook per bit width, stored bit-for-bit in MLX's own packing layout, so the runtime reuses MLX's kernel structure for decode and prefill.

For compatibility with runtimes already being built against it, the on-disk identifiers keep their original names: config.json → jangtq block (version: 2), per-module "mode": "jangtq2", and tensors *.tq2_packed / *.tq2_scales. These are the JANGH format; they are not loadable by, and must not be routed to, JANGTQ v1 loaders.

Quality vs the official FP8 release

15,830 teacher-forced positions on held-out prompts, top-128 renormalized KL.

Bundle Size median KL ↓ mean KL ↓ p90 / p95 / p99 ↓ top-1 ↑ top-5 ↑ top-10 ↑
GLM-5.3-Flash-JANGH2 95.89 GiB 0.0301 0.376 0.98 / 1.97 / 5.24 83.9% 96.8% 98.3%
GLM-5.3-Flash-JANG (affine, previous release) 95.35 GiB 0.0882 0.528 1.50 / 2.55 / 5.67 78.6% 94.6% 96.8%
orcarouter GLM-5.3-Flash-MLX 2bit-lite ¹ 95.4 GiB 0.2122 0.83 — 71.4% 90.8% 94.3%

¹ Measured earlier on the same 20 prompts (15,850 positions), loading its shipped quantized weights natively.

Fidelity vs the bf16 model on real transcripts

Scored against the full bf16 model, run layer by layer from disk. 16 held-out transcripts written by the model itself in its native format: long design and coding reasoning, and multi-step tool conversations with tool results (60,243 assistant positions).

JANGH2 affine JANG (previous)
top-1 agreement with bf16 82.6% 77.1%
median KL vs bf16 0.069 0.158
mean KL vs bf16 0.218 0.356
tool-call points: <tool_call> is the top choice (bf16 agrees on all 41) 40 / 41 35 / 41
median P(<tool_call>) at those points 0.9998 0.968
lowest P(<tool_call>) at those points 0.165 0.245

Agentic fidelity vs the bf16 model itself

Scored against the full bf16 model, run layer by layer from disk. 72 held-out tool-use conversations (25,867 positions; tools and phrasings disjoint from calibration), including 114 "call a tool or answer?" decision points.

JANGH2 affine JANG (previous)
median KL vs bf16 0.579 0.595
top-1 agreement with bf16 57.3% 54.9%
tool-call decisions: same choice as bf16 50 / 50 50 / 50
median P(<tool_call>) at call points (bf16: 0.998) 0.999 0.981
lowest P(<tool_call>) at a call point 0.914 0.784
answer decisions: same next token as bf16 87.5% 50.0%
decisions flipped call ↔ answer vs bf16 0 3

These are synthetic agent transcripts rendered with empty think blocks, so absolute KL is high for every bundle; read the columns against each other. Revision 2 gives up fidelity on this set (revision 1: median KL 0.145, top-1 68.7%) in exchange for ending long reasoning; on assistant output alone top-1 is 92.3% (revision 1: 94.4%). Tool decisions are unchanged. The calibration set includes agent-style conversations from the same generator (different tools and wording), so part of this gain is in-distribution. The FP8 table above is independent of that.

Live behavior (served, temperature 0)

suite JANGH2 affine JANG (previous)
48-case tool-use eval (required / auto × thinking on / off) 48 / 48 48 / 48
16 behavior probes: math, follow-up, tool round trip, efforts low/high/max, image, video 16 / 16 16 / 16
same 16 probes after a server restart (SSD prefix-cache restore) 16 / 16 —

Both bundles saturate these suites, so they check behavior rather than rank the two.

Speed (M5 Max 128 GB, served, interleaved A/B/A/B, fresh prompts, median of 3)

JANGH2 affine JANG (previous)
decode (tok/s) 27.71 / 27.71 26.81 / 26.85
prefill, ~5.2k-token prompt (tok/s) 339 / 361 421 / 416
peak memory while serving 96.7 GiB 96.1 GiB

Two interleaved runs per bundle, same runtime build, each run the median of 3 probes on a never-seen prompt. Decode is ~3% faster than the affine bundle; long-prompt prefill is currently ~15% slower.

What's in the bundle

  • Vision + video: full bf16 vision tower + the consolidated image/video processor config.
  • No MTP: layer 45 is omitted; its bytes went into expert precision.
  • Thinking + agentic:
    • Thinking is ON by default (the template opens <think>).
    • Reasoning efforts are low / high / max (default max). There is no medium: the template renders any other value, and a missing value, as Max.
    • Budget for thinking: the model reasons for 8-10k tokens at low and 17-23k at max on a large design task, and for 4-6k at max even on a small coding task. Use low or high for interactive work and a generous max_tokens.
    • clear_thinking=false preserves thinking in history.
  • Tool calls: GLM's XML dialect (<tool_call>name<arg_key>…</arg_key><arg_value>…</arg_value></tool_call>), declared as tool_parser: glm_xml_args; tool results render as <|observation|>. Hermes-style JSON parsers will not work.
  • Self-describing:
    • config.json carries the JANGH format block (codebook, packing, rotation, method) and a per-module quantization map (126 expert projections, 147 MXFP8, 214 affine 8-bit).
    • jang_config.json records calibration and per-layer expert bits.
    • Raw evaluation results are in evaluation/.
  • Memory: 95.89 GiB weights, 96.7 GiB peak while serving. Fixed-size linear-attention state + compressed-latent KV (~6 KB/token), so long contexts do not balloon memory.
  • Every shard is alignment-safe (zero-copy memory mapping).

Serving contract

  • Sampling: temperature=1.0, top_p=0.95 (vendor defaults), no repetition penalty
  • EOS: [154820, 154827, 154829] · context: 1M native
  • Reasoning: reasoning_effort chat-template kwarg (low / high / max), default max
  • Thinking off: the template always opens <think>. Runtimes must close it in GLM's native form, <think></think> with no whitespace; an R1-style \n</think>\n\n measurably degrades thinking-off tool decisions.

Build details

  • Source: zai-org/GLM-5.3-Flash-BF16 @ a5b45eb
  • Calibration:
    • 600k tokens (web / code / multi-turn chat incl. tool transcripts / math), referenced to the official FP8 release
    • plus a bf16 agentic capture (448 GLM-template tool conversations), with evaluation prompts held out
  • Experts: JANGH (odd-cubic codebook, fp16 per-row scale, Hadamard-32 rotation), GPTQ on all 42 MoE layers with a per-expert importance matrix, bit allocation measured per layer (gate/up 2-3 bit, down 2-3 bit)
  • Revision 2: row scales corrected to unit gain along the source row (jang_config.json → scale_correction)

Quantized and validated by Jinho Jang — eric@jangq.ai

Downloads last month
399
Safetensors
Model size
33B params
Tensor type
U32
·
F32
·
BF16
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JANGQ-AI/GLM-5.3-Flash-JANGH2

Quantized
(155)
this model