vcruz305 commited on
Commit
cc64e77
·
0 Parent(s):

GLM-5.3 MixedK EXL3 3.38 bpw

Browse files
This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. .gitattributes +37 -0
  2. README.md +91 -0
  3. chat_template.jinja +255 -0
  4. config.json +232 -0
  5. generation_config.json +12 -0
  6. model-00001-of-00058.safetensors +3 -0
  7. model-00002-of-00058.safetensors +3 -0
  8. model-00003-of-00058.safetensors +3 -0
  9. model-00004-of-00058.safetensors +3 -0
  10. model-00005-of-00058.safetensors +3 -0
  11. model-00006-of-00058.safetensors +3 -0
  12. model-00007-of-00058.safetensors +3 -0
  13. model-00008-of-00058.safetensors +3 -0
  14. model-00009-of-00058.safetensors +3 -0
  15. model-00010-of-00058.safetensors +3 -0
  16. model-00011-of-00058.safetensors +3 -0
  17. model-00012-of-00058.safetensors +3 -0
  18. model-00013-of-00058.safetensors +3 -0
  19. model-00014-of-00058.safetensors +3 -0
  20. model-00015-of-00058.safetensors +3 -0
  21. model-00016-of-00058.safetensors +3 -0
  22. model-00017-of-00058.safetensors +3 -0
  23. model-00018-of-00058.safetensors +3 -0
  24. model-00019-of-00058.safetensors +3 -0
  25. model-00020-of-00058.safetensors +3 -0
  26. model-00021-of-00058.safetensors +3 -0
  27. model-00022-of-00058.safetensors +3 -0
  28. model-00023-of-00058.safetensors +3 -0
  29. model-00024-of-00058.safetensors +3 -0
  30. model-00025-of-00058.safetensors +3 -0
  31. model-00026-of-00058.safetensors +3 -0
  32. model-00027-of-00058.safetensors +3 -0
  33. model-00028-of-00058.safetensors +3 -0
  34. model-00029-of-00058.safetensors +3 -0
  35. model-00030-of-00058.safetensors +3 -0
  36. model-00031-of-00058.safetensors +3 -0
  37. model-00032-of-00058.safetensors +3 -0
  38. model-00033-of-00058.safetensors +3 -0
  39. model-00034-of-00058.safetensors +3 -0
  40. model-00035-of-00058.safetensors +3 -0
  41. model-00036-of-00058.safetensors +3 -0
  42. model-00037-of-00058.safetensors +3 -0
  43. model-00038-of-00058.safetensors +3 -0
  44. model-00039-of-00058.safetensors +3 -0
  45. model-00040-of-00058.safetensors +3 -0
  46. model-00041-of-00058.safetensors +3 -0
  47. model-00042-of-00058.safetensors +3 -0
  48. model-00043-of-00058.safetensors +3 -0
  49. model-00044-of-00058.safetensors +3 -0
  50. model-00045-of-00058.safetensors +3 -0
.gitattributes ADDED
@@ -0,0 +1,37 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ *.7z filter=lfs diff=lfs merge=lfs -text
2
+ *.arrow filter=lfs diff=lfs merge=lfs -text
3
+ *.bin filter=lfs diff=lfs merge=lfs -text
4
+ *.bz2 filter=lfs diff=lfs merge=lfs -text
5
+ *.ckpt filter=lfs diff=lfs merge=lfs -text
6
+ *.ftz filter=lfs diff=lfs merge=lfs -text
7
+ *.gz filter=lfs diff=lfs merge=lfs -text
8
+ *.h5 filter=lfs diff=lfs merge=lfs -text
9
+ *.joblib filter=lfs diff=lfs merge=lfs -text
10
+ *.lfs.* filter=lfs diff=lfs merge=lfs -text
11
+ *.mlmodel filter=lfs diff=lfs merge=lfs -text
12
+ *.model filter=lfs diff=lfs merge=lfs -text
13
+ *.msgpack filter=lfs diff=lfs merge=lfs -text
14
+ *.npy filter=lfs diff=lfs merge=lfs -text
15
+ *.npz filter=lfs diff=lfs merge=lfs -text
16
+ *.onnx filter=lfs diff=lfs merge=lfs -text
17
+ *.ot filter=lfs diff=lfs merge=lfs -text
18
+ *.parquet filter=lfs diff=lfs merge=lfs -text
19
+ *.pb filter=lfs diff=lfs merge=lfs -text
20
+ *.pickle filter=lfs diff=lfs merge=lfs -text
21
+ *.pkl filter=lfs diff=lfs merge=lfs -text
22
+ *.pt filter=lfs diff=lfs merge=lfs -text
23
+ *.pth filter=lfs diff=lfs merge=lfs -text
24
+ *.rar filter=lfs diff=lfs merge=lfs -text
25
+ *.safetensors filter=lfs diff=lfs merge=lfs -text
26
+ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
27
+ *.tar.* filter=lfs diff=lfs merge=lfs -text
28
+ *.tar filter=lfs diff=lfs merge=lfs -text
29
+ *.tflite filter=lfs diff=lfs merge=lfs -text
30
+ *.tgz filter=lfs diff=lfs merge=lfs -text
31
+ *.wasm filter=lfs diff=lfs merge=lfs -text
32
+ *.xz filter=lfs diff=lfs merge=lfs -text
33
+ *.zip filter=lfs diff=lfs merge=lfs -text
34
+ *.zst filter=lfs diff=lfs merge=lfs -text
35
+ *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
+ model.safetensors.index.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,91 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: zai-org/GLM-5.3
3
+ base_model_relation: quantized
4
+ library_name: exllamav3
5
+ pipeline_tag: text-generation
6
+ license: other
7
+ license_name: glm-5.3
8
+ license_link: https://huggingface.co/zai-org/GLM-5.3
9
+ tags:
10
+ - exl3
11
+ - exllamav3
12
+ - sage
13
+ - mixed-k
14
+ - moe
15
+ - glm
16
+ ---
17
+
18
+ # GLM-5.3 MixedK EXL3 3.38 bpw
19
+
20
+ An EXL3 quantization of [zai-org/GLM-5.3](https://huggingface.co/zai-org/GLM-5.3), made with **SAGE**.
21
+
22
+ SAGE dynamically and intelligently assigns bit widths across the model, making this a MixedK EXL3. It averages 3.38 bits per weight, and every one of the 19,200 routed experts is kept: no pruning and no expert merging.
23
+
24
+ The pack is sized so the weights plus a full 1,048,576-token KV cache at 4 bits fit in the memory of four NVIDIA DGX Sparks.
25
+
26
+ ## Summary
27
+
28
+ | | |
29
+ |---|---|
30
+ | Base model | [zai-org/GLM-5.3](https://huggingface.co/zai-org/GLM-5.3) (78 layers, 256 routed experts per MoE layer, 8 active) |
31
+ | Format | EXL3 |
32
+ | Quantization | SAGE MixedK |
33
+ | Body bitrate | 3.38 bpw nominal, 3.39 bpw including scales |
34
+ | Output head | 8-bit |
35
+ | Routed experts | all 256 per layer kept (19,200 total) |
36
+ | MTP / draft layer | not included |
37
+ | Max context | 1,048,576 tokens (unchanged from the base model) |
38
+ | Total size | 319.0 GB (297.1 GiB), 58 weight shards |
39
+
40
+ ## Quality
41
+
42
+ Measured against the original BF16 model on a held-out evaluation set: text the quantizer never saw.
43
+
44
+ **The held-out set:** 10 sequences of 1,024 tokens each (10,240 scored positions), built from the test splits of public benchmarks:
45
+ - Web text: WikiText-103 (4 sequences)
46
+ - Code: HumanEval (2)
47
+ - Math: GSM8K (2)
48
+ - Chat: UltraChat-200k (2)
49
+
50
+ Each sequence is whole documents packed end to end. Math and chat examples are formatted with GLM-5.3's own chat template. None of this text was used while quantizing the model.
51
+
52
+ **Scoring:** both models read the same tokens. At every position their next-token predictions are compared over the full 154,880-token vocabulary, in float64.
53
+
54
+ | Metric | Result |
55
+ |---|---|
56
+ | Top-1 agreement with BF16 | **92.98%** (9,521 / 10,240) |
57
+ | BF16 top-1 token within the quant's top 5 | 99.38% |
58
+ | Mean KL divergence (BF16 ‖ quant) | 0.0948 |
59
+ | Median KL divergence | 0.0018 |
60
+ | 99th-percentile KL divergence | 1.61 |
61
+ | Top-5 set overlap | 0.827 |
62
+ | Mean NLL, BF16 → quant | 1.0092 → 1.0300 |
63
+ | Perplexity increase | +2.1% |
64
+
65
+ **Scope of these numbers:**
66
+ - The BF16 reference is the original weights run through the same ExLlamaV3 runtime, not the vendor's own implementation.
67
+ - Scores come from full-sequence forward passes at 1,024 tokens. Long-context and cached generation were not part of this evaluation.
68
+
69
+ ## Memory budget: four DGX Sparks at 1M context
70
+
71
+ | Item | Size |
72
+ |---|---|
73
+ | Weights | 319.0 GB |
74
+ | KV cache, 1,048,576 tokens at Q4 (MLA latent plus indexer keys) | ~39.7 GB |
75
+ | Total | ~358.7 GB of 512 GB (4 × 128 GB) |
76
+
77
+ The rest is left for activations, the runtime and the OS. This is a sizing target only: four-node serving and full 1M-token inference have not been validated yet.
78
+
79
+ ## Runtime
80
+
81
+ GLM-5.3 support (DSA sparse attention, the F32 router bias and MixedK experts) is in a fork of [exllamav3](https://github.com/turboderp-org/exllamav3), at commit `affc194d5476710f167f30729e2508e577759612`. Stock exllamav3 cannot load this model yet. Installation and usage instructions will be added here when the runtime is released.
82
+
83
+ ## Download
84
+
85
+ ```bash
86
+ hf download vcruz305/GLM-5.3-EXL3-3.38bpw --local-dir GLM-5.3-EXL3-3.38bpw
87
+ ```
88
+
89
+ ## License
90
+
91
+ Same license as the base model; see [zai-org/GLM-5.3](https://huggingface.co/zai-org/GLM-5.3).
chat_template.jinja ADDED
@@ -0,0 +1,255 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [gMASK]<sop>
2
+ {%- set effective_reasoning_effort = reasoning_effort if reasoning_effort is defined and reasoning_effort in ['low', 'high'] else 'max' -%}
3
+ {%- if effective_reasoning_effort is not none -%}<|system|>Reasoning Effort: {{ effective_reasoning_effort | capitalize }}{%- endif -%}
4
+ {%- set clear_thinking = clear_thinking if clear_thinking is defined else false -%}
5
+ {%- if tools -%}
6
+ {%- macro tool_to_json(tool) -%}
7
+ {%- set ns_tool = namespace(first=true) -%}
8
+ {{ '{' -}}
9
+ {%- for k, v in tool.items() -%}
10
+ {%- if k != 'defer_loading' and k != 'strict' -%}
11
+ {%- if not ns_tool.first -%}{{- ', ' -}}{%- endif -%}
12
+ {%- set ns_tool.first = false -%}
13
+ "{{ k }}": {{ v | tojson(ensure_ascii=False) }}
14
+ {%- endif -%}
15
+ {%- endfor -%}
16
+ {{- '}' -}}
17
+ {%- endmacro -%}
18
+ {%- macro tool_references_to_response(refs) -%}
19
+ {{- '<tool_response><tools>\n' -}}
20
+ {%- for tr in refs -%}
21
+ {%- for tool in tools -%}
22
+ {%- if 'function' in tool -%}
23
+ {%- set tool = tool['function'] -%}
24
+ {%- endif -%}
25
+ {%- if tool.name == tr.name -%}
26
+ {{- tool_to_json(tool) + '\n' -}}
27
+ {%- endif -%}
28
+ {%- endfor -%}
29
+ {%- endfor -%}
30
+ {{- '</tools></tool_response>' -}}
31
+ {%- endmacro -%}
32
+ <|system|>
33
+ # Tools
34
+
35
+ You may call one or more functions to assist with the user query.
36
+
37
+ You are provided with function signatures within <tools></tools> XML tags:
38
+ <tools>
39
+ {% for tool in tools %}
40
+ {%- if 'function' in tool -%}
41
+ {%- set tool = tool['function'] -%}
42
+ {%- endif -%}
43
+ {% if tool.defer_loading is not defined or not tool.defer_loading %}
44
+ {{ tool_to_json(tool) }}
45
+ {% endif %}
46
+ {% endfor %}
47
+ </tools>
48
+
49
+ For each function call, output the function name and arguments within the following XML format:
50
+ <tool_call>{function-name}<arg_key>{arg-key-1}</arg_key><arg_value>{arg-value-1}</arg_value><arg_key>{arg-key-2}</arg_key><arg_value>{arg-value-2}</arg_value>...</tool_call>{%- endif -%}
51
+ {%- macro visible_text(content) -%}
52
+ {%- if content is string -%}
53
+ {{- content }}
54
+ {%- elif content is iterable and content is not mapping -%}
55
+ {%- for item in content -%}
56
+ {%- if item is mapping and item.type == 'text' -%}
57
+ {{- item.text }}
58
+ {%- elif item is string -%}
59
+ {{- item }}
60
+ {%- elif item is mapping and item.type in ['image', 'image_url', 'video', 'video_url', 'audio', 'audio_url', 'input_audio'] -%}
61
+ {%- set media_type = item.type | replace('_url', '') | replace('input_', '') -%}
62
+ {{- "<reminder>You are unable to process this " ~ media_type ~ " because you don't have multi-modal input ability. Try different methods.</reminder>" }}
63
+ {%- endif -%}
64
+ {%- endfor -%}
65
+ {%- elif content is not none -%}
66
+ {{- content }}
67
+ {%- endif -%}
68
+ {%- endmacro -%}
69
+ {%- macro tool_response(text) -%}
70
+ {{- '<tool_response>' + text + '</tool_response>' -}}
71
+ {%- endmacro -%}
72
+ {%- macro render_tool_response(m) -%}
73
+ {%- if m.content is string -%}
74
+ {{- tool_response(m.content) -}}
75
+ {%- elif m.content and m.content is not mapping and m.content.0.type == "tool_reference" -%}
76
+ {{- tool_references_to_response(m.content) -}}
77
+ {%- elif is_list_of_outputs(m) -%}
78
+ {%- for tr in m.content -%}
79
+ {%- if tr.output is iterable and tr.output is not string and tr.output is not mapping and tr.output and tr.output.0.type == "tool_reference" -%}
80
+ {{- tool_references_to_response(tr.output) -}}
81
+ {%- else -%}
82
+ {{- tool_response(visible_text(tr.output)) -}}
83
+ {%- endif -%}
84
+ {%- endfor -%}
85
+ {%- else -%}
86
+ {{- tool_response(visible_text(m.content)) -}}
87
+ {%- endif -%}
88
+ {%- endmacro -%}
89
+ {%- macro id_of(obj) -%}
90
+ {%- if obj.tool_call_id -%}
91
+ {{- obj.tool_call_id -}}
92
+ {%- elif obj.id -%}
93
+ {{- obj.id -}}
94
+ {%- endif -%}
95
+ {%- endmacro -%}
96
+ {%- macro is_list_of_outputs(m) -%}
97
+ {%- if m.content and m.content.0.output is defined -%}1{%- endif -%}
98
+ {%- endmacro -%}
99
+ {%- macro has_dup_tool_result_id(lo, hi, target) -%}
100
+ {%- set ns_cnt = namespace(n=0) -%}
101
+ {%- for k in range(lo, hi + 1) -%}
102
+ {%- set m = messages[k] -%}
103
+ {%- if is_list_of_outputs(m) -%}
104
+ {%- for entry in m.content -%}
105
+ {%- if id_of(entry) == target -%}
106
+ {%- set ns_cnt.n = ns_cnt.n + 1 -%}
107
+ {%- endif -%}
108
+ {%- endfor -%}
109
+ {%- elif id_of(m) == target -%}
110
+ {%- set ns_cnt.n = ns_cnt.n + 1 -%}
111
+ {%- endif -%}
112
+ {%- if ns_cnt.n > 1 -%}{%- break -%}{%- endif -%}
113
+ {%- endfor -%}
114
+ {%- if ns_cnt.n > 1 -%}1{%- endif -%}
115
+ {%- endmacro -%}
116
+ {%- macro tc_id_exists(tcs, target) -%}
117
+ {%- set ns_f = namespace(found=false) -%}
118
+ {%- for tc in tcs -%}
119
+ {%- if id_of(tc) == target -%}
120
+ {%- set ns_f.found = true -%}
121
+ {%- break -%}
122
+ {%- endif -%}
123
+ {%- endfor -%}
124
+ {%- if ns_f.found -%}1{%- endif -%}
125
+ {%- endmacro -%}
126
+ {%- set ns = namespace(last_user_index=-1) -%}
127
+ {%- for m in messages %}
128
+ {%- if m.role == 'user' %}
129
+ {%- set ns.last_user_index = loop.index0 -%}
130
+ {%- endif %}
131
+ {%- endfor %}
132
+ {%- for m in messages -%}
133
+ {%- if m.role == 'user' -%}<|user|>{{ visible_text(m.content) }}
134
+ {%- elif m.role == 'assistant' -%}
135
+ <|assistant|>
136
+ {%- set content = visible_text(m.content) %}
137
+ {%- if m.reasoning_content is string %}
138
+ {%- set reasoning_content = m.reasoning_content %}
139
+ {%- elif '</think>' in content %}
140
+ {%- set reasoning_content = content.split('</think>')[0].split('<think>')[-1] %}
141
+ {%- set content = content.split('</think>')[-1] %}
142
+ {%- endif %}
143
+ {%- if (not clear_thinking or loop.index0 > ns.last_user_index) and reasoning_content is defined -%}
144
+ {{ '<think>' + reasoning_content + '</think>'}}
145
+ {%- else -%}
146
+ {{ '<think></think>' }}
147
+ {%- endif -%}
148
+ {%- if content.strip() -%}
149
+ {{ content.strip() }}
150
+ {%- endif -%}
151
+ {% if m.tool_calls %}
152
+ {% for tc in m.tool_calls %}
153
+ {%- if tc.function %}
154
+ {%- set tc = tc.function %}
155
+ {%- endif %}
156
+ {{- '<tool_call>' + tc.name -}}
157
+ {% set _args = tc.arguments %}{% for k, v in _args.items() %}<arg_key>{{ k }}</arg_key><arg_value>{{ v | tojson(ensure_ascii=False) if v is not string else v }}</arg_value>{% endfor %}</tool_call>{% endfor %}
158
+ {% endif %}
159
+ {%- elif m.role == 'tool' -%}
160
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
161
+ {{- '<|observation|>' -}}
162
+ {%- set block_start = loop.index0 -%}
163
+ {%- set ns_blk = namespace(end=block_start) -%}
164
+ {%- for j in range(block_start, messages|length) -%}
165
+ {%- if messages[j].role == 'tool' -%}
166
+ {%- set ns_blk.end = j -%}
167
+ {%- else -%}
168
+ {%- break -%}
169
+ {%- endif -%}
170
+ {%- endfor -%}
171
+ {%- set ns_a = namespace(tool_calls=none) -%}
172
+ {%- if block_start > 0 and messages[block_start - 1].role == 'assistant' and messages[block_start - 1].tool_calls -%}
173
+ {%- set ns_a.tool_calls = messages[block_start - 1].tool_calls -%}
174
+ {%- endif -%}
175
+ {%- set ns_chk = namespace(can_sort=true) -%}
176
+ {%- if not ns_a.tool_calls -%}
177
+ {%- set ns_chk.can_sort = false -%}
178
+ {%- else -%}
179
+ {%- for k in range(block_start, ns_blk.end + 1) -%}
180
+ {%- if not ns_chk.can_sort -%}{%- break -%}{%- endif -%}
181
+ {%- set m = messages[k] -%}
182
+ {%- if is_list_of_outputs(m) -%}
183
+ {%- for entry in m.content -%}
184
+ {%- if not ns_chk.can_sort -%}{%- break -%}{%- endif -%}
185
+ {%- set eid = id_of(entry) -%}
186
+ {%- if not eid -%}
187
+ {%- set ns_chk.can_sort = false -%}
188
+ {%- elif has_dup_tool_result_id(block_start, ns_blk.end, eid) -%}
189
+ {%- set ns_chk.can_sort = false -%}
190
+ {%- elif not tc_id_exists(ns_a.tool_calls, eid) -%}
191
+ {%- set ns_chk.can_sort = false -%}
192
+ {%- endif -%}
193
+ {%- endfor -%}
194
+ {%- else -%}
195
+ {%- set tk_id = id_of(m) -%}
196
+ {%- if not tk_id -%}
197
+ {%- set ns_chk.can_sort = false -%}
198
+ {%- elif has_dup_tool_result_id(block_start, ns_blk.end, tk_id) -%}
199
+ {%- set ns_chk.can_sort = false -%}
200
+ {%- elif not tc_id_exists(ns_a.tool_calls, tk_id) -%}
201
+ {%- set ns_chk.can_sort = false -%}
202
+ {%- endif -%}
203
+ {%- endif -%}
204
+ {%- endfor -%}
205
+ {%- for i in range(ns_a.tool_calls | length) -%}
206
+ {%- if not ns_chk.can_sort -%}{%- break -%}{%- endif -%}
207
+ {%- set tc_id = id_of(ns_a.tool_calls[i]) -%}
208
+ {%- if not tc_id -%}
209
+ {%- set ns_chk.can_sort = false -%}
210
+ {%- endif -%}
211
+ {%- for j in range(i + 1, ns_a.tool_calls | length) -%}
212
+ {%- if id_of(ns_a.tool_calls[j]) == tc_id -%}
213
+ {%- set ns_chk.can_sort = false -%}
214
+ {%- break -%}
215
+ {%- endif -%}
216
+ {%- endfor -%}
217
+ {%- endfor -%}
218
+ {%- endif -%}
219
+ {%- if ns_chk.can_sort -%}
220
+ {%- for tc in ns_a.tool_calls -%}
221
+ {%- set tc_id = id_of(tc) -%}
222
+ {%- for k in range(block_start, ns_blk.end + 1) -%}
223
+ {%- set m = messages[k] -%}
224
+ {%- if is_list_of_outputs(m) -%}
225
+ {%- for entry in m.content -%}
226
+ {%- set eid = id_of(entry) -%}
227
+ {%- if eid == tc_id -%}
228
+ {%- if entry.output is iterable and entry.output is not string and entry.output is not mapping and entry.output and entry.output.0.type == "tool_reference" -%}
229
+ {{- tool_references_to_response(entry.output) -}}
230
+ {%- else -%}
231
+ {{- tool_response(visible_text(entry.output)) -}}
232
+ {%- endif -%}
233
+ {%- endif -%}
234
+ {%- endfor -%}
235
+ {%- else -%}
236
+ {%- set tk_id = id_of(m) -%}
237
+ {%- if tk_id == tc_id -%}
238
+ {{- render_tool_response(m) -}}
239
+ {%- endif -%}
240
+ {%- endif -%}
241
+ {%- endfor -%}
242
+ {%- endfor -%}
243
+ {%- else -%}
244
+ {%- for k in range(block_start, ns_blk.end + 1) -%}
245
+ {{- render_tool_response(messages[k]) -}}
246
+ {%- endfor -%}
247
+ {%- endif -%}
248
+ {% endif -%}
249
+ {%- elif m.role == 'system' -%}
250
+ <|system|>{{ visible_text(m.content) }}
251
+ {%- endif -%}
252
+ {%- endfor -%}
253
+ {%- if add_generation_prompt -%}
254
+ <|assistant|>{{- '<think>' -}}
255
+ {%- endif -%}
config.json ADDED
@@ -0,0 +1,232 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "GlmMoeDsaForCausalLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "dtype": "bfloat16",
8
+ "eos_token_id": [
9
+ 154820,
10
+ 154827,
11
+ 154829
12
+ ],
13
+ "ep_size": 1,
14
+ "first_k_dense_replace": 3,
15
+ "head_dim": 192,
16
+ "hidden_act": "silu",
17
+ "hidden_size": 6144,
18
+ "index_head_dim": 128,
19
+ "index_n_heads": 32,
20
+ "index_share_for_mtp_iteration": true,
21
+ "index_skip_topk_offset": 3,
22
+ "index_topk": 2048,
23
+ "index_topk_freq": 4,
24
+ "index_topk_pattern": null,
25
+ "indexer_rope_interleave": true,
26
+ "indexer_types": [
27
+ "full",
28
+ "full",
29
+ "full",
30
+ "shared",
31
+ "shared",
32
+ "shared",
33
+ "full",
34
+ "shared",
35
+ "shared",
36
+ "shared",
37
+ "full",
38
+ "shared",
39
+ "shared",
40
+ "shared",
41
+ "full",
42
+ "shared",
43
+ "shared",
44
+ "shared",
45
+ "full",
46
+ "shared",
47
+ "shared",
48
+ "shared",
49
+ "full",
50
+ "shared",
51
+ "shared",
52
+ "shared",
53
+ "full",
54
+ "shared",
55
+ "shared",
56
+ "shared",
57
+ "full",
58
+ "shared",
59
+ "shared",
60
+ "shared",
61
+ "full",
62
+ "shared",
63
+ "shared",
64
+ "shared",
65
+ "full",
66
+ "shared",
67
+ "shared",
68
+ "shared",
69
+ "full",
70
+ "shared",
71
+ "shared",
72
+ "shared",
73
+ "full",
74
+ "shared",
75
+ "shared",
76
+ "shared",
77
+ "full",
78
+ "shared",
79
+ "shared",
80
+ "shared",
81
+ "full",
82
+ "shared",
83
+ "shared",
84
+ "shared",
85
+ "full",
86
+ "shared",
87
+ "shared",
88
+ "shared",
89
+ "full",
90
+ "shared",
91
+ "shared",
92
+ "shared",
93
+ "full",
94
+ "shared",
95
+ "shared",
96
+ "shared",
97
+ "full",
98
+ "shared",
99
+ "shared",
100
+ "shared",
101
+ "full",
102
+ "shared",
103
+ "shared",
104
+ "shared"
105
+ ],
106
+ "initializer_range": 0.02,
107
+ "intermediate_size": 12288,
108
+ "kv_lora_rank": 512,
109
+ "max_position_embeddings": 1048576,
110
+ "mlp_layer_types": [
111
+ "dense",
112
+ "dense",
113
+ "dense",
114
+ "sparse",
115
+ "sparse",
116
+ "sparse",
117
+ "sparse",
118
+ "sparse",
119
+ "sparse",
120
+ "sparse",
121
+ "sparse",
122
+ "sparse",
123
+ "sparse",
124
+ "sparse",
125
+ "sparse",
126
+ "sparse",
127
+ "sparse",
128
+ "sparse",
129
+ "sparse",
130
+ "sparse",
131
+ "sparse",
132
+ "sparse",
133
+ "sparse",
134
+ "sparse",
135
+ "sparse",
136
+ "sparse",
137
+ "sparse",
138
+ "sparse",
139
+ "sparse",
140
+ "sparse",
141
+ "sparse",
142
+ "sparse",
143
+ "sparse",
144
+ "sparse",
145
+ "sparse",
146
+ "sparse",
147
+ "sparse",
148
+ "sparse",
149
+ "sparse",
150
+ "sparse",
151
+ "sparse",
152
+ "sparse",
153
+ "sparse",
154
+ "sparse",
155
+ "sparse",
156
+ "sparse",
157
+ "sparse",
158
+ "sparse",
159
+ "sparse",
160
+ "sparse",
161
+ "sparse",
162
+ "sparse",
163
+ "sparse",
164
+ "sparse",
165
+ "sparse",
166
+ "sparse",
167
+ "sparse",
168
+ "sparse",
169
+ "sparse",
170
+ "sparse",
171
+ "sparse",
172
+ "sparse",
173
+ "sparse",
174
+ "sparse",
175
+ "sparse",
176
+ "sparse",
177
+ "sparse",
178
+ "sparse",
179
+ "sparse",
180
+ "sparse",
181
+ "sparse",
182
+ "sparse",
183
+ "sparse",
184
+ "sparse",
185
+ "sparse",
186
+ "sparse",
187
+ "sparse",
188
+ "sparse"
189
+ ],
190
+ "model_type": "glm_moe_dsa",
191
+ "moe_intermediate_size": 2048,
192
+ "moe_layer_freq": 1,
193
+ "moe_router_dtype": "float32",
194
+ "n_group": 1,
195
+ "n_routed_experts": 256,
196
+ "n_shared_experts": 1,
197
+ "norm_topk_prob": true,
198
+ "num_attention_heads": 64,
199
+ "num_experts_per_tok": 8,
200
+ "num_hidden_layers": 78,
201
+ "num_key_value_heads": 64,
202
+ "num_nextn_predict_layers": 0,
203
+ "pad_token_id": 154820,
204
+ "pretraining_tp": 1,
205
+ "q_lora_rank": 2048,
206
+ "qk_head_dim": 256,
207
+ "qk_nope_head_dim": 192,
208
+ "qk_rope_head_dim": 64,
209
+ "rms_norm_eps": 1e-05,
210
+ "rope_interleave": true,
211
+ "rope_parameters": {
212
+ "rope_theta": 8000000,
213
+ "rope_type": "default"
214
+ },
215
+ "routed_scaling_factor": 2.5,
216
+ "scoring_func": "sigmoid",
217
+ "tie_word_embeddings": false,
218
+ "topk_group": 1,
219
+ "topk_method": "noaux_tc",
220
+ "transformers_version": "5.15.0",
221
+ "use_cache": true,
222
+ "v_head_dim": 256,
223
+ "vocab_size": 154880,
224
+ "quantization_config": {
225
+ "quant_method": "exl3",
226
+ "version": "1.4.9",
227
+ "bits": 3,
228
+ "head_bits": 8,
229
+ "out_scales": "always",
230
+ "codebook": "mul1"
231
+ }
232
+ }
generation_config.json ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "eos_token_id": [
4
+ 154820,
5
+ 154827,
6
+ 154829
7
+ ],
8
+ "pad_token_id": 154820,
9
+ "temperature": 1.0,
10
+ "top_p": 0.95,
11
+ "transformers_version": "5.12.0"
12
+ }
model-00001-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:94feb26e0d4e00d743c3984d1c23aa5aab9dcb89e1741561c122f19835368a56
3
+ size 6196618489
model-00002-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4613b1f1c6a3c255e03d61b2714cb59051330694351e32848e451b5a331ed8a2
3
+ size 7511872122
model-00003-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5df5c573c9cbea092331c9f31809fd7b08333117983cb6654595b0adb89ab877
3
+ size 7536920622
model-00004-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:511b6d9704943dc473e197736a2f3c6ea49cd578f375134a156fa5208a0f47ce
3
+ size 7573213826
model-00005-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e7aca28a4e1514dccf9bdb518dfde007127ebe4128eeff0fbb0b49b88ff5c333
3
+ size 7577821332
model-00006-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cd32deb3666e7871ad3f34a1d765acea332c8e02256527a56f8493fd458e2db9
3
+ size 7545957056
model-00007-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2cdeb7921d43e9f675e2bc2d81ff78c0f9e930ebbfd254619ef65cec9df72088
3
+ size 7539548308
model-00008-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ba719b88ed6fe6f42d8ed0337fe32f533c86716b371aad1993c4c486ffc16223
3
+ size 7505586872
model-00009-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e7e53706c148b3f552201c34178146e856c218a6805eb40579e803e96690f5d5
3
+ size 7511761012
model-00010-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2ec676f9acffc1918c67d29a764937b226e13166a6e92c7831250c87edffb01e
3
+ size 7494052536
model-00011-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7daf94316da01fb2675b8e4fd9c39bb8199baf5fa2fe14c886bc09c57608f189
3
+ size 7514906740
model-00012-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:08b1b55961942457407db8c9a2a67898cb94008995b9da69b35e315f370fe274
3
+ size 7616211672
model-00013-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9a627c2de90ba661bdce4b6e75955fcc685e8684ac04d8ec32cbd2eb6d3e43d8
3
+ size 7547936901
model-00014-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d0a09dc1c7de0e931926f6e4090e6375c1d5b5fadcbb14e622600332bd9306ca
3
+ size 7534422712
model-00015-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:81b41e069829d2a0d7a5bafdf858a116e488e8a08f659ce45eceda00764677bb
3
+ size 7525916789
model-00016-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3bbdc7dd2364c6ffbf9c225afd30ef832fcedbd47def9338a9e29419622cdf6f
3
+ size 7514991288
model-00017-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:dc8be30abaa840ab43c7555535ac3d99cc4986160b42cd4fc85ef07943953a70
3
+ size 7958388917
model-00018-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:16d020096ffa8e457f3b9d89ba82b2fb4507a35a32257733fba5a2546de8b0bf
3
+ size 8557243336
model-00019-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d73eacb466b05115fe97f446a62e419b1ca1373908d89387b50372afa2735bc4
3
+ size 4471965853
model-00020-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8149a2c2cadf9e6628f9916aadcf5339bdd80cf05769a68cf6d71561fc9d3c53
3
+ size 4489908976
model-00021-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0fbbada4ad9dd7b433a8ec3e9af6626d10626b03e09a605bb74a09328cf38f40
3
+ size 4432237280
model-00022-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7d94d3e18c715dec89e15c450cf9302061db917356048be992a7e7c16188926a
3
+ size 4415984344
model-00023-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6be59ae4dc37d1fc852d9ff3af9a431088c260f95992abf3aed2c000a5566402
3
+ size 4445227133
model-00024-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e3257b11337a5ee6f5fa9bd87545b67851839e9188df9f376955fa9acfba6e64
3
+ size 4532376312
model-00025-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c5f5330b36cf620da4fdb1e531b724047eb0ee5e342b96e7ca82088c2846768d
3
+ size 4548104952
model-00026-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6a4f6caa109bbae2ecc009525a7a1f85dfb24f7a82d858bd0eb143c0d9ca2649
3
+ size 4552823552
model-00027-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bcfbba7fe63d20c65d538249fb31b3801bb8bdc8f4e313e35d5f65dc6b0d68e9
3
+ size 4560046269
model-00028-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fb595b4afacf2576e500baf5439048e35b041adfb69110cbfcf1a1d564446219
3
+ size 4554396408
model-00029-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a9f3fa791aee78437e15ea7aebe921d454a6d42c3dafc6655b205b9aa4498bed
3
+ size 4556493568
model-00030-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3259c2a5814215520bf3ce99c254bc3616fc6f39ab98954a3c83e2107be1cd5a
3
+ size 4555969280
model-00031-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:acd8d4402f07b18f2edc9c36bf847ad97e78096b17f196a9e66f1e6c9ea1d66b
3
+ size 4571056317
model-00032-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a26572707ac1536adf0c25af52669c9de3dbc8cdf4299a447d593193bfe7398c
3
+ size 4563833600
model-00033-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7cfd140b88381effab4aec238e1f6f860fcf481162b6334772b2a69d643fb24a
3
+ size 4560687872
model-00034-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:83668c92d2ed270c2980588bd7663b9eb6c1a3aa7783bb7d99e8649062d2e835
3
+ size 4573270784
model-00035-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:142dca1e40e8b4cf15c96c97f0b98023532668ff9897fbe8cdd3b3ac9bf01a00
3
+ size 4571580589
model-00036-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:73ddbb114d224e72ba9d42d4c422ec42cc88156b341ab3def6fb87e0b79492c3
3
+ size 4576416520
model-00037-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e793d162b82da1cc8d67fd7754e23f9a5465a7d51d2b483ed70209740efe49ea
3
+ size 4568552192
model-00038-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:588775295e5cbaa6f99b19f9b6c10064c97b49713c3578f968e1dda183c64036
3
+ size 4575367936
model-00039-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:82ec3a493d7a9949ab8f61b35c37353f3b223a555438a88b815fd6f78088553e
3
+ size 4574202053
model-00040-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e7cdee4402f02ac67ceb157f90bc477a347faa883a61a35754c02ad0b4b08164
3
+ size 4574843648
model-00041-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a4a93f936b42b03a5e7fd6f2c2e8354a52eb6fca6bf4a039d171811b311f4afb
3
+ size 4582707976
model-00042-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bb4be22d4ef9317a09bfff5985870544f9752f3c8a74d0ad3bb218789741d3e8
3
+ size 4581135112
model-00043-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8bd4e234aed228483bfd8679134fc12b0493a630b6d651b17171bb38779c1f8b
3
+ size 4586784965
model-00044-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3f46bbeb3e8f4cb857e988cc76284e2f43ebb50a28cf3b0cb0c7fde2c7af45af
3
+ size 4586377992
model-00045-of-00058.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4362d1ead3b5b00d2482ddeeadefc8310b962b9a52751be51ba321a2b8ba69fd
3
+ size 4595290888