--- language: - en library_name: mlx license: mit pipeline_tag: image-text-to-text base_model: zai-org/GLM-5.3-Flash tags: - mlx - jang - jangh - quantized - apple-silicon - vision - video - reasoning - thinking - agent - tool-use - glm5_next - moe - gptq - imatrix ---

MLX Studio / vMLX — run JANG models on Apple Silicon

Run JANG models in MLX Studio / vMLX

JANGQ-AI

> ✅ **Runtime: supported in vMLX (Python).** Use **vMLX 1.6.68 or newer** (native JANGH loading; the vMLX 1.6.71 > release was validated with GLM JANGH on native reasoning efforts and chat tool/media continuations). Older vMLX builds refuse the > bundle at load time rather than producing wrong output. > Known limitation (vMLX 1.6.71 notes): video perception is imperfect — the model can describe overlay or > transition text that is not in the clip. # JANGQ-AI/GLM-5.3-Flash-JANGH2 **GLM-5.3-Flash for 128 GB Macs.** Same size as our previous affine release, with **2.9x lower median KL**, **+5.3 points top-1**, and faster decode. A JANGH bundle of [zai-org/GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash): a 300B-class MoE (288 routed experts, top-8 + shared expert) with KDA linear attention, sparse attention, and vision + video towers. - **Routed experts**: JANGH at 2-3 bits. Codebook-quantized with per-row scales and a blockwise Hadamard rotation, calibrated with GPTQ + a per-expert importance matrix. Bits are placed per layer by measurement. - **Everything else**: 8-bit (MXFP8 where the source weights sit on an FP8-MX grid, affine 8-bit elsewhere) or full precision. - **Vision tower**: kept in bf16. ## Revision 2: long reasoning now ends on its own The first upload could not end long reasoning. After a few thousand reasoning tokens the probability of `` was hundreds to thousands of times too low, so the model kept thinking, and the text eventually degraded into repetition. With the default effort (`max`) that range is reached on ordinary tasks. Cause: error-minimizing row scales shrink every low-bit matrix a little (12% at 2 bits). Three matrices per expert and 42 MoE layers deep, that weakens exactly the strong, late decisions, and "stop thinking" is one of them. Revision 2 changes **only the row scales of the routed experts** (126 small tensors). The quantized codes, the size, the format and the speed are unchanged, and no runtime change is needed. | where the reference model ends its reasoning | bf16 | **revision 2** | revision 1 | affine JANG (previous) | |---|---|---|---|---| | P(``), held-out design task, after 8,025 reasoning tokens | 0.82 | **0.75** | 0.03 | 0.02 | | P(``), coding task, after 3,288 reasoning tokens | 0.99 | **0.94** | 0.19 | - | | long design prompt, served, vendor sampling: reasoning ended by the model | **revision 2** | revision 1 | |---|---|---| | effort `low` | **3 / 3** | 0 / 5 | | effort `high` | **2 / 2** | - | | effort `max` (40k-token budget; the reference model needs 17-23k reasoning tokens here) | **1 / 1** | 0 / 2 | Your copy is revision 2 if `jang_config.json` has a `scale_correction` entry. This replaces `JANGQ-AI/GLM-5.3-Flash-JANG` and `-JANG-MTP` (and was briefly published as `GLM-5.3-Flash-JANGTQ2`). ## JANGH vs JANGTQ JANGH replaces our earlier JANGTQ (TurboQuant) expert format. **The TurboQuant rotation is no longer used:** no seeded random-sign rotation and no per-dimension Beta-distribution codebook, so runtimes do not need to reproduce any RNG or dimension-specific tables. JANGH uses a fixed blockwise Hadamard-32 transform (no random signs) and a two-constant codebook per bit width, stored bit-for-bit in MLX's own packing layout, so the runtime reuses MLX's kernel structure for decode and prefill. For compatibility with runtimes already being built against it, the on-disk identifiers keep their original names: `config.json` → `jangtq` block (`version: 2`), per-module `"mode": "jangtq2"`, and tensors `*.tq2_packed` / `*.tq2_scales`. These are the JANGH format; they are not loadable by, and must not be routed to, JANGTQ v1 loaders. ## Quality vs the official FP8 release 15,830 teacher-forced positions on held-out prompts, top-128 renormalized KL. | Bundle | Size | median KL ↓ | mean KL ↓ | p90 / p95 / p99 ↓ | top-1 ↑ | top-5 ↑ | top-10 ↑ | |---|---|---|---|---|---|---|---| | **GLM-5.3-Flash-JANGH2** | **95.89 GiB** | **0.0301** | **0.376** | **0.98 / 1.97 / 5.24** | **83.9%** | **96.8%** | **98.3%** | | GLM-5.3-Flash-JANG (affine, previous release) | 95.35 GiB | 0.0882 | 0.528 | 1.50 / 2.55 / 5.67 | 78.6% | 94.6% | 96.8% | | orcarouter GLM-5.3-Flash-MLX `2bit-lite` ¹ | 95.4 GiB | 0.2122 | 0.83 | — | 71.4% | 90.8% | 94.3% | ¹ Measured earlier on the same 20 prompts (15,850 positions), loading its shipped quantized weights natively. ## Fidelity vs the bf16 model on real transcripts Scored against the full bf16 model, run layer by layer from disk. 16 held-out transcripts written by the model itself in its native format: long design and coding reasoning, and multi-step tool conversations with tool results (60,243 assistant positions). | | **JANGH2** | affine JANG (previous) | |---|---|---| | top-1 agreement with bf16 | **82.6%** | 77.1% | | median KL vs bf16 | **0.069** | 0.158 | | mean KL vs bf16 | **0.218** | 0.356 | | tool-call points: `` is the top choice (bf16 agrees on all 41) | **40 / 41** | 35 / 41 | | median P(``) at those points | **0.9998** | 0.968 | | lowest P(``) at those points | 0.165 | **0.245** | ## Agentic fidelity vs the bf16 model itself Scored against the full bf16 model, run layer by layer from disk. 72 held-out tool-use conversations (25,867 positions; tools and phrasings disjoint from calibration), including 114 "call a tool or answer?" decision points. | | **JANGH2** | affine JANG (previous) | |---|---|---| | median KL vs bf16 | **0.579** | 0.595 | | top-1 agreement with bf16 | **57.3%** | 54.9% | | tool-call decisions: same choice as bf16 | **50 / 50** | 50 / 50 | | median P(``) at call points (bf16: 0.998) | **0.999** | 0.981 | | lowest P(``) at a call point | **0.914** | 0.784 | | answer decisions: same next token as bf16 | **87.5%** | 50.0% | | decisions flipped call ↔ answer vs bf16 | **0** | 3 | These are synthetic agent transcripts rendered with empty think blocks, so absolute KL is high for every bundle; read the columns against each other. Revision 2 gives up fidelity on this set (revision 1: median KL 0.145, top-1 68.7%) in exchange for ending long reasoning; on assistant output alone top-1 is 92.3% (revision 1: 94.4%). Tool decisions are unchanged. The calibration set includes agent-style conversations from the same generator (different tools and wording), so part of this gain is in-distribution. The FP8 table above is independent of that. ## Live behavior (served, temperature 0) | suite | **JANGH2** | affine JANG (previous) | |---|---|---| | 48-case tool-use eval (required / auto × thinking on / off) | **48 / 48** | 48 / 48 | | 16 behavior probes: math, follow-up, tool round trip, efforts low/high/max, image, video | **16 / 16** | 16 / 16 | | same 16 probes after a server restart (SSD prefix-cache restore) | **16 / 16** | — | Both bundles saturate these suites, so they check behavior rather than rank the two. ## Speed (M5 Max 128 GB, served, interleaved A/B/A/B, fresh prompts, median of 3) | | **JANGH2** | affine JANG (previous) | |---|---|---| | decode (tok/s) | **27.71 / 27.71** | 26.81 / 26.85 | | prefill, ~5.2k-token prompt (tok/s) | 339 / 361 | 421 / 416 | | peak memory while serving | 96.7 GiB | 96.1 GiB | Two interleaved runs per bundle, same runtime build, each run the median of 3 probes on a never-seen prompt. Decode is **~3% faster** than the affine bundle; long-prompt prefill is currently ~15% slower. ## What's in the bundle - **Vision + video**: full bf16 vision tower + the consolidated image/video processor config. - **No MTP**: layer 45 is omitted; its bytes went into expert precision. - **Thinking + agentic**: - Thinking is ON by default (the template opens ``). - Reasoning efforts are **`low` / `high` / `max`** (default `max`). There is no `medium`: the template renders any other value, and a missing value, as Max. - Budget for thinking: the model reasons for 8-10k tokens at `low` and 17-23k at `max` on a large design task, and for 4-6k at `max` even on a small coding task. Use `low` or `high` for interactive work and a generous `max_tokens`. - `clear_thinking=false` preserves thinking in history. - **Tool calls**: GLM's XML dialect (`name……`), declared as `tool_parser: glm_xml_args`; tool results render as `<|observation|>`. Hermes-style JSON parsers will not work. - **Self-describing**: - `config.json` carries the JANGH format block (codebook, packing, rotation, method) and a per-module `quantization` map (126 expert projections, 147 MXFP8, 214 affine 8-bit). - `jang_config.json` records calibration and per-layer expert bits. - Raw evaluation results are in `evaluation/`. - **Memory**: 95.89 GiB weights, 96.7 GiB peak while serving. Fixed-size linear-attention state + compressed-latent KV (~6 KB/token), so long contexts do not balloon memory. - Every shard is alignment-safe (zero-copy memory mapping). ## Serving contract - Sampling: `temperature=1.0, top_p=0.95` (vendor defaults), no repetition penalty - EOS: `[154820, 154827, 154829]` · context: 1M native - Reasoning: `reasoning_effort` chat-template kwarg (`low` / `high` / `max`), default `max` - Thinking off: the template always opens ``. Runtimes must close it in GLM's native form, `` with no whitespace; an R1-style `\n\n\n` measurably degrades thinking-off tool decisions. ## Build details - Source: `zai-org/GLM-5.3-Flash-BF16` @ `a5b45eb` - Calibration: - 600k tokens (web / code / multi-turn chat incl. tool transcripts / math), referenced to the official FP8 release - plus a bf16 agentic capture (448 GLM-template tool conversations), with evaluation prompts held out - Experts: JANGH (odd-cubic codebook, fp16 per-row scale, Hadamard-32 rotation), GPTQ on all 42 MoE layers with a per-expert importance matrix, bit allocation measured per layer (gate/up 2-3 bit, down 2-3 bit) - Revision 2: row scales corrected to unit gain along the source row (`jang_config.json` → `scale_correction`) Quantized and validated by **Jinho Jang** — eric@jangq.ai