---
language:
- en
library_name: mlx
license: mit
pipeline_tag: image-text-to-text
base_model: zai-org/GLM-5.3-Flash
tags:
- mlx
- jang
- jangh
- quantized
- apple-silicon
- vision
- video
- reasoning
- thinking
- agent
- tool-use
- glm5_next
- moe
- gptq
- imatrix
---

> ✅ **Runtime: supported in vMLX (Python).** Use **vMLX 1.6.68 or newer** (native JANGH loading; the vMLX 1.6.71
> release was validated with GLM JANGH on native reasoning efforts and chat tool/media continuations). Older vMLX builds refuse the
> bundle at load time rather than producing wrong output.
> Known limitation (vMLX 1.6.71 notes): video perception is imperfect — the model can describe overlay or
> transition text that is not in the clip.
# JANGQ-AI/GLM-5.3-Flash-JANGH2
**GLM-5.3-Flash for 128 GB Macs.** Same size as our previous affine release, with **2.9x lower median KL**, **+5.3
points top-1**, and faster decode.
A JANGH bundle of [zai-org/GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash): a 300B-class MoE (288
routed experts, top-8 + shared expert) with KDA linear attention, sparse attention, and vision + video towers.
- **Routed experts**: JANGH at 2-3 bits. Codebook-quantized with per-row scales and a blockwise Hadamard rotation,
calibrated with GPTQ + a per-expert importance matrix. Bits are placed per layer by measurement.
- **Everything else**: 8-bit (MXFP8 where the source weights sit on an FP8-MX grid, affine 8-bit elsewhere) or full
precision.
- **Vision tower**: kept in bf16.
## Revision 2: long reasoning now ends on its own
The first upload could not end long reasoning. After a few thousand reasoning tokens the probability of `` was
hundreds to thousands of times too low, so the model kept thinking, and the text eventually degraded into repetition.
With the default effort (`max`) that range is reached on ordinary tasks.
Cause: error-minimizing row scales shrink every low-bit matrix a little (12% at 2 bits). Three matrices per expert
and 42 MoE layers deep, that weakens exactly the strong, late decisions, and "stop thinking" is one of them.
Revision 2 changes **only the row scales of the routed experts** (126 small tensors). The quantized codes, the size,
the format and the speed are unchanged, and no runtime change is needed.
| where the reference model ends its reasoning | bf16 | **revision 2** | revision 1 | affine JANG (previous) |
|---|---|---|---|---|
| P(``), held-out design task, after 8,025 reasoning tokens | 0.82 | **0.75** | 0.03 | 0.02 |
| P(``), coding task, after 3,288 reasoning tokens | 0.99 | **0.94** | 0.19 | - |
| long design prompt, served, vendor sampling: reasoning ended by the model | **revision 2** | revision 1 |
|---|---|---|
| effort `low` | **3 / 3** | 0 / 5 |
| effort `high` | **2 / 2** | - |
| effort `max` (40k-token budget; the reference model needs 17-23k reasoning tokens here) | **1 / 1** | 0 / 2 |
Your copy is revision 2 if `jang_config.json` has a `scale_correction` entry.
This replaces `JANGQ-AI/GLM-5.3-Flash-JANG` and `-JANG-MTP` (and was briefly published as `GLM-5.3-Flash-JANGTQ2`).
## JANGH vs JANGTQ
JANGH replaces our earlier JANGTQ (TurboQuant) expert format. **The TurboQuant rotation is no longer used:** no
seeded random-sign rotation and no per-dimension Beta-distribution codebook, so runtimes do not need to reproduce any RNG
or dimension-specific tables. JANGH uses a fixed blockwise Hadamard-32 transform (no random signs) and a two-constant
codebook per bit width, stored bit-for-bit in MLX's own packing layout, so the runtime reuses MLX's kernel structure for
decode and prefill.
For compatibility with runtimes already being built against it, the on-disk identifiers keep their original names:
`config.json` → `jangtq` block (`version: 2`), per-module `"mode": "jangtq2"`, and tensors `*.tq2_packed` /
`*.tq2_scales`. These are the JANGH format; they are not loadable by, and must not be routed to, JANGTQ v1 loaders.
## Quality vs the official FP8 release
15,830 teacher-forced positions on held-out prompts, top-128 renormalized KL.
| Bundle | Size | median KL ↓ | mean KL ↓ | p90 / p95 / p99 ↓ | top-1 ↑ | top-5 ↑ | top-10 ↑ |
|---|---|---|---|---|---|---|---|
| **GLM-5.3-Flash-JANGH2** | **95.89 GiB** | **0.0301** | **0.376** | **0.98 / 1.97 / 5.24** | **83.9%** | **96.8%** | **98.3%** |
| GLM-5.3-Flash-JANG (affine, previous release) | 95.35 GiB | 0.0882 | 0.528 | 1.50 / 2.55 / 5.67 | 78.6% | 94.6% | 96.8% |
| orcarouter GLM-5.3-Flash-MLX `2bit-lite` ¹ | 95.4 GiB | 0.2122 | 0.83 | — | 71.4% | 90.8% | 94.3% |
¹ Measured earlier on the same 20 prompts (15,850 positions), loading its shipped quantized weights natively.
## Fidelity vs the bf16 model on real transcripts
Scored against the full bf16 model, run layer by layer from disk. 16 held-out transcripts written by the model
itself in its native format: long design and coding reasoning, and multi-step tool conversations with tool results
(60,243 assistant positions).
| | **JANGH2** | affine JANG (previous) |
|---|---|---|
| top-1 agreement with bf16 | **82.6%** | 77.1% |
| median KL vs bf16 | **0.069** | 0.158 |
| mean KL vs bf16 | **0.218** | 0.356 |
| tool-call points: `` is the top choice (bf16 agrees on all 41) | **40 / 41** | 35 / 41 |
| median P(``) at those points | **0.9998** | 0.968 |
| lowest P(``) at those points | 0.165 | **0.245** |
## Agentic fidelity vs the bf16 model itself
Scored against the full bf16 model, run layer by layer from disk. 72 held-out tool-use conversations (25,867
positions; tools and phrasings disjoint from calibration), including 114 "call a tool or answer?" decision points.
| | **JANGH2** | affine JANG (previous) |
|---|---|---|
| median KL vs bf16 | **0.579** | 0.595 |
| top-1 agreement with bf16 | **57.3%** | 54.9% |
| tool-call decisions: same choice as bf16 | **50 / 50** | 50 / 50 |
| median P(``) at call points (bf16: 0.998) | **0.999** | 0.981 |
| lowest P(``) at a call point | **0.914** | 0.784 |
| answer decisions: same next token as bf16 | **87.5%** | 50.0% |
| decisions flipped call ↔ answer vs bf16 | **0** | 3 |
These are synthetic agent transcripts rendered with empty think blocks, so absolute KL is high for every bundle; read
the columns against each other. Revision 2 gives up fidelity on this set (revision 1: median KL 0.145, top-1 68.7%) in
exchange for ending long reasoning; on assistant output alone top-1 is 92.3% (revision 1: 94.4%). Tool decisions are
unchanged.
The calibration set includes agent-style conversations from the same generator (different tools and wording), so
part of this gain is in-distribution. The FP8 table above is independent of that.
## Live behavior (served, temperature 0)
| suite | **JANGH2** | affine JANG (previous) |
|---|---|---|
| 48-case tool-use eval (required / auto × thinking on / off) | **48 / 48** | 48 / 48 |
| 16 behavior probes: math, follow-up, tool round trip, efforts low/high/max, image, video | **16 / 16** | 16 / 16 |
| same 16 probes after a server restart (SSD prefix-cache restore) | **16 / 16** | — |
Both bundles saturate these suites, so they check behavior rather than rank the two.
## Speed (M5 Max 128 GB, served, interleaved A/B/A/B, fresh prompts, median of 3)
| | **JANGH2** | affine JANG (previous) |
|---|---|---|
| decode (tok/s) | **27.71 / 27.71** | 26.81 / 26.85 |
| prefill, ~5.2k-token prompt (tok/s) | 339 / 361 | 421 / 416 |
| peak memory while serving | 96.7 GiB | 96.1 GiB |
Two interleaved runs per bundle, same runtime build, each run the median of 3 probes on a never-seen prompt.
Decode is **~3% faster** than the affine bundle; long-prompt prefill is currently ~15% slower.
## What's in the bundle
- **Vision + video**: full bf16 vision tower + the consolidated image/video processor config.
- **No MTP**: layer 45 is omitted; its bytes went into expert precision.
- **Thinking + agentic**:
- Thinking is ON by default (the template opens ``).
- Reasoning efforts are **`low` / `high` / `max`** (default `max`). There is no `medium`: the template renders any
other value, and a missing value, as Max.
- Budget for thinking: the model reasons for 8-10k tokens at `low` and 17-23k at `max` on a large design task, and
for 4-6k at `max` even on a small coding task. Use `low` or `high` for interactive work and a generous `max_tokens`.
- `clear_thinking=false` preserves thinking in history.
- **Tool calls**: GLM's XML dialect (`name……`),
declared as `tool_parser: glm_xml_args`; tool results render as `<|observation|>`. Hermes-style JSON parsers will
not work.
- **Self-describing**:
- `config.json` carries the JANGH format block (codebook, packing, rotation, method) and a per-module
`quantization` map (126 expert projections, 147 MXFP8, 214 affine 8-bit).
- `jang_config.json` records calibration and per-layer expert bits.
- Raw evaluation results are in `evaluation/`.
- **Memory**: 95.89 GiB weights, 96.7 GiB peak while serving. Fixed-size linear-attention state +
compressed-latent KV (~6 KB/token), so long contexts do not balloon memory.
- Every shard is alignment-safe (zero-copy memory mapping).
## Serving contract
- Sampling: `temperature=1.0, top_p=0.95` (vendor defaults), no repetition penalty
- EOS: `[154820, 154827, 154829]` · context: 1M native
- Reasoning: `reasoning_effort` chat-template kwarg (`low` / `high` / `max`), default `max`
- Thinking off: the template always opens ``. Runtimes must close it in GLM's native form, ``
with no whitespace; an R1-style `\n\n\n` measurably degrades thinking-off tool decisions.
## Build details
- Source: `zai-org/GLM-5.3-Flash-BF16` @ `a5b45eb`
- Calibration:
- 600k tokens (web / code / multi-turn chat incl. tool transcripts / math), referenced to the official FP8 release
- plus a bf16 agentic capture (448 GLM-template tool conversations), with evaluation prompts held out
- Experts: JANGH (odd-cubic codebook, fp16 per-row scale, Hadamard-32 rotation), GPTQ on all 42 MoE layers with a
per-expert importance matrix, bit allocation measured per layer (gate/up 2-3 bit, down 2-3 bit)
- Revision 2: row scales corrected to unit gain along the source row (`jang_config.json` → `scale_correction`)
Quantized and validated by **Jinho Jang** — eric@jangq.ai