jangq commited on
Commit
9e37a8f
·
verified ·
1 Parent(s): 5b23997

Add model card

Browse files
Files changed (1) hide show
  1. README.md +137 -0
README.md ADDED
@@ -0,0 +1,137 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ library_name: mlx
5
+ license: mit
6
+ pipeline_tag: image-text-to-text
7
+ base_model: zai-org/GLM-5.3-Flash
8
+ tags:
9
+ - mlx
10
+ - jang
11
+ - jangtq
12
+ - quantized
13
+ - apple-silicon
14
+ - vision
15
+ - video
16
+ - reasoning
17
+ - thinking
18
+ - agent
19
+ - tool-use
20
+ - glm5_next
21
+ - moe
22
+ - gptq
23
+ - imatrix
24
+ ---
25
+
26
+ <p align="center">
27
+ <img src="./jangq-logo.png" alt="JANGQ" width="220">
28
+ &nbsp;&nbsp;&nbsp;
29
+ <img src="./vmlx-logo.png" alt="vMLX" width="90">
30
+ </p>
31
+
32
+ > ⚠️ **Runtime not released yet.** This bundle uses **JANGTQ v2**, a new routed-expert format, on a new
33
+ > architecture (`glm5_next`: KDA linear attention + MLA/DSA hybrid + mHC). No released vMLX or Osaurus build can load
34
+ > it today. They refuse it at load time rather than producing wrong output. Access is gated until runtime support
35
+ > ships. The numbers below were measured with the internal vMLX build this bundle was made for.
36
+
37
+ # JANGQ-AI/GLM-5.3-Flash-JANGTQ2
38
+
39
+ **GLM-5.3-Flash for 128 GB Macs.** Same size as our previous affine release, with **2.7x lower median KL**, **+5.1
40
+ points top-1**, and faster decode.
41
+
42
+ A JANGTQ v2 bundle of [zai-org/GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash): a 300B-class MoE (288
43
+ routed experts, top-8 + shared expert) with KDA linear attention, sparse attention, and vision + video towers.
44
+ - **Routed experts**: JANGTQ v2 at 2-3 bits. Codebook-quantized with per-row scales and a blockwise Hadamard rotation,
45
+ calibrated with GPTQ + a per-expert importance matrix. Bits are placed per layer by measurement.
46
+ - **Everything else**: 8-bit (MXFP8 where the source weights sit on an FP8-MX grid, affine 8-bit elsewhere) or full
47
+ precision.
48
+ - **Vision tower**: kept in bf16.
49
+
50
+ This replaces `JANGQ-AI/GLM-5.3-Flash-JANG` and `-JANG-MTP`.
51
+
52
+ ## Quality vs the official FP8 release
53
+ 15,830 teacher-forced positions on held-out prompts, top-128 renormalized KL.
54
+
55
+ | Bundle | Size | median KL ↓ | mean KL ↓ | p90 / p95 / p99 ↓ | top-1 ↑ | top-5 ↑ | top-10 ↑ |
56
+ |---|---|---|---|---|---|---|---|
57
+ | **GLM-5.3-Flash-JANGTQ2** | **95.89 GiB** | **0.0323** | **0.351** | **0.89 / 1.76 / 4.53** | **83.7%** | **97.0%** | **98.5%** |
58
+ | GLM-5.3-Flash-JANG (affine, previous release) | 95.35 GiB | 0.0882 | 0.528 | 1.50 / 2.55 / 5.67 | 78.6% | 94.6% | 96.8% |
59
+ | orcarouter GLM-5.3-Flash-MLX `2bit-lite` ¹ | 95.4 GiB | 0.2122 | 0.83 | — | 71.4% | 90.8% | 94.3% |
60
+
61
+ ¹ Measured earlier on the same 20 prompts (15,850 positions), loading its shipped quantized weights natively.
62
+
63
+ ## Agentic fidelity vs the bf16 model itself
64
+ Scored against the full bf16 model, run layer by layer from disk. 72 held-out tool-use conversations (25,867
65
+ positions; tools and phrasings disjoint from calibration), including 114 "call a tool or answer?" decision points.
66
+
67
+ | | **JANGTQ2** | affine JANG (previous) |
68
+ |---|---|---|
69
+ | median KL vs bf16 | **0.145** | 0.595 |
70
+ | top-1 agreement with bf16 | **68.7%** | 54.9% |
71
+ | tool-call decisions: same choice as bf16 | **50 / 50** | 50 / 50 |
72
+ | median P(`<tool_call>`) at call points (bf16: 0.998) | **0.999** | 0.981 |
73
+ | lowest P(`<tool_call>`) at a call point | **0.987** | 0.784 |
74
+ | answer decisions: same next token as bf16 | **89.1%** | 50.0% |
75
+ | decisions flipped call ↔ answer vs bf16 | **0** | 3 |
76
+
77
+ These are synthetic agent transcripts, so absolute KL is high for every bundle; read the columns against each other.
78
+ The calibration set includes agent-style conversations from the same generator (different tools and wording), so
79
+ part of this gain is in-distribution. The FP8 table above is independent of that.
80
+
81
+ ## Live behavior (served, temperature 0)
82
+
83
+ | suite | **JANGTQ2** | affine JANG (previous) |
84
+ |---|---|---|
85
+ | 48-case tool-use eval (required / auto × thinking on / off) | **48 / 48** | 48 / 48 |
86
+ | 16 behavior probes: math, follow-up, tool round trip, efforts low/high/max, image, video | **16 / 16** | 16 / 16 |
87
+ | same 16 probes after a server restart (SSD prefix-cache restore) | **16 / 16** | — |
88
+
89
+ Both bundles saturate these suites, so they check behavior rather than rank the two.
90
+
91
+ ## Speed (M5 Max 128 GB, served, interleaved A/B/A/B, fresh prompts, median of 3)
92
+
93
+ | | **JANGTQ2** | affine JANG (previous) |
94
+ |---|---|---|
95
+ | decode (tok/s) | **27.87 / 27.88** | 26.81 / 26.85 |
96
+ | prefill, ~5.2k-token prompt (tok/s) | 348 / 390 | 421 / 416 |
97
+ | peak memory while serving | 96.7 GiB | 96.1 GiB |
98
+
99
+ Two interleaved runs per bundle, same runtime build, each run the median of 3 probes on a never-seen prompt.
100
+ Decode is **~4% faster** than the affine bundle; long-prompt prefill is currently ~12% slower.
101
+
102
+ ## What's in the bundle
103
+ - **Vision + video**: full bf16 vision tower + the consolidated image/video processor config.
104
+ - **No MTP**: layer 45 is omitted; its bytes went into expert precision.
105
+ - **Thinking + agentic**:
106
+ - Thinking is ON by default (the template opens `<think>`).
107
+ - Reasoning efforts are **`low` / `high` / `max`** (default `max`). There is no `medium`: the template renders any
108
+ other value as Max.
109
+ - `clear_thinking=false` preserves thinking in history.
110
+ - **Tool calls**: GLM's XML dialect (`<tool_call>name<arg_key>…</arg_key><arg_value>…</arg_value></tool_call>`),
111
+ declared as `tool_parser: glm_xml_args`; tool results render as `<|observation|>`. Hermes-style JSON parsers will
112
+ not work.
113
+ - **Self-describing**:
114
+ - `config.json` carries the JANGTQ v2 format block (codebook, packing, rotation, method) and a per-module
115
+ `quantization` map (126 expert projections, 147 MXFP8, 214 affine 8-bit).
116
+ - `jang_config.json` records calibration and per-layer expert bits.
117
+ - Raw evaluation results are in `evaluation/`.
118
+ - **Memory**: 95.89 GiB weights, 96.7 GiB peak while serving. Fixed-size linear-attention state +
119
+ compressed-latent KV (~6 KB/token), so long contexts do not balloon memory.
120
+ - Every shard is alignment-safe (zero-copy memory mapping).
121
+
122
+ ## Serving contract
123
+ - Sampling: `temperature=1.0, top_p=0.95` (vendor defaults), no repetition penalty
124
+ - EOS: `[154820, 154827, 154829]` · context: 1M native
125
+ - Reasoning: `reasoning_effort` chat-template kwarg (`low` / `high` / `max`), default `max`
126
+ - Thinking off: the template always opens `<think>`. Runtimes must close it in GLM's native form, `<think></think>`
127
+ with no whitespace; an R1-style `\n</think>\n\n` measurably degrades thinking-off tool decisions.
128
+
129
+ ## Build details
130
+ - Source: `zai-org/GLM-5.3-Flash-BF16` @ `a5b45eb`
131
+ - Calibration:
132
+ - 600k tokens (web / code / multi-turn chat incl. tool transcripts / math), referenced to the official FP8 release
133
+ - plus a bf16 agentic capture (448 GLM-template tool conversations), with evaluation prompts held out
134
+ - Experts: JANGTQ v2 (odd-cubic codebook, fp16 per-row scale, Hadamard-32 rotation), GPTQ on all 42 MoE layers with a
135
+ per-expert importance matrix, bit allocation measured per layer (gate/up 2-3 bit, down 2-3 bit)
136
+
137
+ Quantized and validated by **Jinho Jang** — eric@jangq.ai