Image-Text-to-Text
MLX
Safetensors
English
Chinese
glm5_next
jang
janght
jangt
jangh
quantized
apple-silicon
vision
video
reasoning
thinking
agent
tool-use
speculative-decoding
dflash
Mixture of Experts
conversational
8-bit precision
Instructions to use JANGQ-AI/GLM-5.3-Flash-JANGHT2.4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use JANGQ-AI/GLM-5.3-Flash-JANGHT2.4 with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("JANGQ-AI/GLM-5.3-Flash-JANGHT2.4") config = load_config("JANGQ-AI/GLM-5.3-Flash-JANGHT2.4") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use JANGQ-AI/GLM-5.3-Flash-JANGHT2.4 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "JANGQ-AI/GLM-5.3-Flash-JANGHT2.4"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "JANGQ-AI/GLM-5.3-Flash-JANGHT2.4" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use JANGQ-AI/GLM-5.3-Flash-JANGHT2.4 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "JANGQ-AI/GLM-5.3-Flash-JANGHT2.4"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default JANGQ-AI/GLM-5.3-Flash-JANGHT2.4
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use JANGQ-AI/GLM-5.3-Flash-JANGHT2.4 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "JANGQ-AI/GLM-5.3-Flash-JANGHT2.4"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "JANGQ-AI/GLM-5.3-Flash-JANGHT2.4" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Add model card
Browse files
README.md
ADDED
|
@@ -0,0 +1,137 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language:
|
| 3 |
+
- en
|
| 4 |
+
library_name: mlx
|
| 5 |
+
license: mit
|
| 6 |
+
pipeline_tag: image-text-to-text
|
| 7 |
+
base_model: zai-org/GLM-5.3-Flash
|
| 8 |
+
tags:
|
| 9 |
+
- mlx
|
| 10 |
+
- jang
|
| 11 |
+
- jangtq
|
| 12 |
+
- quantized
|
| 13 |
+
- apple-silicon
|
| 14 |
+
- vision
|
| 15 |
+
- video
|
| 16 |
+
- reasoning
|
| 17 |
+
- thinking
|
| 18 |
+
- agent
|
| 19 |
+
- tool-use
|
| 20 |
+
- glm5_next
|
| 21 |
+
- moe
|
| 22 |
+
- gptq
|
| 23 |
+
- imatrix
|
| 24 |
+
---
|
| 25 |
+
|
| 26 |
+
<p align="center">
|
| 27 |
+
<img src="./jangq-logo.png" alt="JANGQ" width="220">
|
| 28 |
+
|
| 29 |
+
<img src="./vmlx-logo.png" alt="vMLX" width="90">
|
| 30 |
+
</p>
|
| 31 |
+
|
| 32 |
+
> ⚠️ **Runtime not released yet.** This bundle uses **JANGTQ v2**, a new routed-expert format, on a new
|
| 33 |
+
> architecture (`glm5_next`: KDA linear attention + MLA/DSA hybrid + mHC). No released vMLX or Osaurus build can load
|
| 34 |
+
> it today. They refuse it at load time rather than producing wrong output. Access is gated until runtime support
|
| 35 |
+
> ships. The numbers below were measured with the internal vMLX build this bundle was made for.
|
| 36 |
+
|
| 37 |
+
# JANGQ-AI/GLM-5.3-Flash-JANGTQ2
|
| 38 |
+
|
| 39 |
+
**GLM-5.3-Flash for 128 GB Macs.** Same size as our previous affine release, with **2.7x lower median KL**, **+5.1
|
| 40 |
+
points top-1**, and faster decode.
|
| 41 |
+
|
| 42 |
+
A JANGTQ v2 bundle of [zai-org/GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash): a 300B-class MoE (288
|
| 43 |
+
routed experts, top-8 + shared expert) with KDA linear attention, sparse attention, and vision + video towers.
|
| 44 |
+
- **Routed experts**: JANGTQ v2 at 2-3 bits. Codebook-quantized with per-row scales and a blockwise Hadamard rotation,
|
| 45 |
+
calibrated with GPTQ + a per-expert importance matrix. Bits are placed per layer by measurement.
|
| 46 |
+
- **Everything else**: 8-bit (MXFP8 where the source weights sit on an FP8-MX grid, affine 8-bit elsewhere) or full
|
| 47 |
+
precision.
|
| 48 |
+
- **Vision tower**: kept in bf16.
|
| 49 |
+
|
| 50 |
+
This replaces `JANGQ-AI/GLM-5.3-Flash-JANG` and `-JANG-MTP`.
|
| 51 |
+
|
| 52 |
+
## Quality vs the official FP8 release
|
| 53 |
+
15,830 teacher-forced positions on held-out prompts, top-128 renormalized KL.
|
| 54 |
+
|
| 55 |
+
| Bundle | Size | median KL ↓ | mean KL ↓ | p90 / p95 / p99 ↓ | top-1 ↑ | top-5 ↑ | top-10 ↑ |
|
| 56 |
+
|---|---|---|---|---|---|---|---|
|
| 57 |
+
| **GLM-5.3-Flash-JANGTQ2** | **95.89 GiB** | **0.0323** | **0.351** | **0.89 / 1.76 / 4.53** | **83.7%** | **97.0%** | **98.5%** |
|
| 58 |
+
| GLM-5.3-Flash-JANG (affine, previous release) | 95.35 GiB | 0.0882 | 0.528 | 1.50 / 2.55 / 5.67 | 78.6% | 94.6% | 96.8% |
|
| 59 |
+
| orcarouter GLM-5.3-Flash-MLX `2bit-lite` ¹ | 95.4 GiB | 0.2122 | 0.83 | — | 71.4% | 90.8% | 94.3% |
|
| 60 |
+
|
| 61 |
+
¹ Measured earlier on the same 20 prompts (15,850 positions), loading its shipped quantized weights natively.
|
| 62 |
+
|
| 63 |
+
## Agentic fidelity vs the bf16 model itself
|
| 64 |
+
Scored against the full bf16 model, run layer by layer from disk. 72 held-out tool-use conversations (25,867
|
| 65 |
+
positions; tools and phrasings disjoint from calibration), including 114 "call a tool or answer?" decision points.
|
| 66 |
+
|
| 67 |
+
| | **JANGTQ2** | affine JANG (previous) |
|
| 68 |
+
|---|---|---|
|
| 69 |
+
| median KL vs bf16 | **0.145** | 0.595 |
|
| 70 |
+
| top-1 agreement with bf16 | **68.7%** | 54.9% |
|
| 71 |
+
| tool-call decisions: same choice as bf16 | **50 / 50** | 50 / 50 |
|
| 72 |
+
| median P(`<tool_call>`) at call points (bf16: 0.998) | **0.999** | 0.981 |
|
| 73 |
+
| lowest P(`<tool_call>`) at a call point | **0.987** | 0.784 |
|
| 74 |
+
| answer decisions: same next token as bf16 | **89.1%** | 50.0% |
|
| 75 |
+
| decisions flipped call ↔ answer vs bf16 | **0** | 3 |
|
| 76 |
+
|
| 77 |
+
These are synthetic agent transcripts, so absolute KL is high for every bundle; read the columns against each other.
|
| 78 |
+
The calibration set includes agent-style conversations from the same generator (different tools and wording), so
|
| 79 |
+
part of this gain is in-distribution. The FP8 table above is independent of that.
|
| 80 |
+
|
| 81 |
+
## Live behavior (served, temperature 0)
|
| 82 |
+
|
| 83 |
+
| suite | **JANGTQ2** | affine JANG (previous) |
|
| 84 |
+
|---|---|---|
|
| 85 |
+
| 48-case tool-use eval (required / auto × thinking on / off) | **48 / 48** | 48 / 48 |
|
| 86 |
+
| 16 behavior probes: math, follow-up, tool round trip, efforts low/high/max, image, video | **16 / 16** | 16 / 16 |
|
| 87 |
+
| same 16 probes after a server restart (SSD prefix-cache restore) | **16 / 16** | — |
|
| 88 |
+
|
| 89 |
+
Both bundles saturate these suites, so they check behavior rather than rank the two.
|
| 90 |
+
|
| 91 |
+
## Speed (M5 Max 128 GB, served, interleaved A/B/A/B, fresh prompts, median of 3)
|
| 92 |
+
|
| 93 |
+
| | **JANGTQ2** | affine JANG (previous) |
|
| 94 |
+
|---|---|---|
|
| 95 |
+
| decode (tok/s) | **27.87 / 27.88** | 26.81 / 26.85 |
|
| 96 |
+
| prefill, ~5.2k-token prompt (tok/s) | 348 / 390 | 421 / 416 |
|
| 97 |
+
| peak memory while serving | 96.7 GiB | 96.1 GiB |
|
| 98 |
+
|
| 99 |
+
Two interleaved runs per bundle, same runtime build, each run the median of 3 probes on a never-seen prompt.
|
| 100 |
+
Decode is **~4% faster** than the affine bundle; long-prompt prefill is currently ~12% slower.
|
| 101 |
+
|
| 102 |
+
## What's in the bundle
|
| 103 |
+
- **Vision + video**: full bf16 vision tower + the consolidated image/video processor config.
|
| 104 |
+
- **No MTP**: layer 45 is omitted; its bytes went into expert precision.
|
| 105 |
+
- **Thinking + agentic**:
|
| 106 |
+
- Thinking is ON by default (the template opens `<think>`).
|
| 107 |
+
- Reasoning efforts are **`low` / `high` / `max`** (default `max`). There is no `medium`: the template renders any
|
| 108 |
+
other value as Max.
|
| 109 |
+
- `clear_thinking=false` preserves thinking in history.
|
| 110 |
+
- **Tool calls**: GLM's XML dialect (`<tool_call>name<arg_key>…</arg_key><arg_value>…</arg_value></tool_call>`),
|
| 111 |
+
declared as `tool_parser: glm_xml_args`; tool results render as `<|observation|>`. Hermes-style JSON parsers will
|
| 112 |
+
not work.
|
| 113 |
+
- **Self-describing**:
|
| 114 |
+
- `config.json` carries the JANGTQ v2 format block (codebook, packing, rotation, method) and a per-module
|
| 115 |
+
`quantization` map (126 expert projections, 147 MXFP8, 214 affine 8-bit).
|
| 116 |
+
- `jang_config.json` records calibration and per-layer expert bits.
|
| 117 |
+
- Raw evaluation results are in `evaluation/`.
|
| 118 |
+
- **Memory**: 95.89 GiB weights, 96.7 GiB peak while serving. Fixed-size linear-attention state +
|
| 119 |
+
compressed-latent KV (~6 KB/token), so long contexts do not balloon memory.
|
| 120 |
+
- Every shard is alignment-safe (zero-copy memory mapping).
|
| 121 |
+
|
| 122 |
+
## Serving contract
|
| 123 |
+
- Sampling: `temperature=1.0, top_p=0.95` (vendor defaults), no repetition penalty
|
| 124 |
+
- EOS: `[154820, 154827, 154829]` · context: 1M native
|
| 125 |
+
- Reasoning: `reasoning_effort` chat-template kwarg (`low` / `high` / `max`), default `max`
|
| 126 |
+
- Thinking off: the template always opens `<think>`. Runtimes must close it in GLM's native form, `<think></think>`
|
| 127 |
+
with no whitespace; an R1-style `\n</think>\n\n` measurably degrades thinking-off tool decisions.
|
| 128 |
+
|
| 129 |
+
## Build details
|
| 130 |
+
- Source: `zai-org/GLM-5.3-Flash-BF16` @ `a5b45eb`
|
| 131 |
+
- Calibration:
|
| 132 |
+
- 600k tokens (web / code / multi-turn chat incl. tool transcripts / math), referenced to the official FP8 release
|
| 133 |
+
- plus a bf16 agentic capture (448 GLM-template tool conversations), with evaluation prompts held out
|
| 134 |
+
- Experts: JANGTQ v2 (odd-cubic codebook, fp16 per-row scale, Hadamard-32 rotation), GPTQ on all 42 MoE layers with a
|
| 135 |
+
per-expert importance matrix, bit allocation measured per layer (gate/up 2-3 bit, down 2-3 bit)
|
| 136 |
+
|
| 137 |
+
Quantized and validated by **Jinho Jang** — eric@jangq.ai
|