infosave commited on
Commit
5f7b65c
·
verified ·
1 Parent(s): 6ff7f8e

Model card: Hy-MT2 1.8B/7B/30B-A3B in CMF (cortiq 0.7.0)

Browse files
Files changed (1) hide show
  1. README.md +288 -0
README.md ADDED
@@ -0,0 +1,288 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: cortiq
4
+ base_model:
5
+ - tencent/Hy-MT2-1.8B
6
+ - tencent/Hy-MT2-7B
7
+ - tencent/Hy-MT2-30B-A3B
8
+ base_model_relation: quantized
9
+ pipeline_tag: translation
10
+ tags:
11
+ - cmf
12
+ - cortiq
13
+ - quantized
14
+ - q4tp
15
+ - q1t
16
+ - ternary
17
+ - translation
18
+ - moe
19
+ - hunyuan
20
+ - 4-bit
21
+ language:
22
+ - zh
23
+ - en
24
+ - fr
25
+ - pt
26
+ - es
27
+ - ja
28
+ - tr
29
+ - ru
30
+ - ar
31
+ - ko
32
+ - th
33
+ - it
34
+ - de
35
+ - vi
36
+ - ms
37
+ - id
38
+ - tl
39
+ - hi
40
+ - pl
41
+ - cs
42
+ - nl
43
+ - km
44
+ - my
45
+ - fa
46
+ - gu
47
+ - ur
48
+ - te
49
+ - mr
50
+ - he
51
+ - bn
52
+ - ta
53
+ - uk
54
+ - bo
55
+ - kk
56
+ - mn
57
+ - ug
58
+ ---
59
+
60
+ # Tencent Hy-MT2 — CMF (1.8B · 7B · 30B-A3B)
61
+
62
+ Tencent's [Hy-MT2](https://huggingface.co/collections/tencent/hy-mt2)
63
+ translation family — 33 languages, instruction-following translation
64
+ (terminology, style, placeholders, structured data) — as
65
+ [CMF](https://github.com/infosave2007/cmf) files for `cortiq`: one
66
+ memory-mapped file per model with the weights, the vendor tokenizer, the
67
+ exact chat template and per-tensor hashes. The same file runs on the CPU,
68
+ on Vulkan (NVIDIA / AMD / Intel), DX12 and Apple Metal; no Python, no
69
+ PyTorch, no CUDA toolkit.
70
+
71
+ ```bash
72
+ cargo install cortiq-cli --version 0.7.0 # first release with HunYuan support
73
+
74
+ hf download infosave/Hy-MT2-cmf hy-mt2-7b-q4tp.cmf --local-dir .
75
+ cortiq verify hy-mt2-7b-q4tp.cmf
76
+ cortiq run hy-mt2-7b-q4tp.cmf --greedy --prompt "Translate the following text into English. Note that you should only output the translated result without any additional explanation:
77
+
78
+ Сегодня хорошая погода, и мы пойдём гулять в парк."
79
+ ```
80
+
81
+ **Requires cortiq 0.7.0 or newer.** Earlier runtimes do not know the
82
+ `hunyuan_v1_dense` / `hy_v3` architectures (and 0.6.9 crashes on any
83
+ 4-bit embedding — fixed in the same release).
84
+
85
+ ## Files
86
+
87
+ | model | file | source | bits | size | wikitext-2 ppl¹ | MT ppl² |
88
+ |---|---|---|---|---:|---:|---:|
89
+ | Hy-MT2-1.8B | `hy-mt2-1.8b-q4tp.cmf` | bf16 checkpoint | 4.17 | 950.5 MB | 27.13 | 14.02 |
90
+ | Hy-MT2-1.8B | `hy-mt2-1.8b-q1t.cmf` | Tencent's 1.25-bit STQ (AngelSlim) | 2.25 | 694.9 MB | 4 460³ | 27.88 |
91
+ | Hy-MT2-7B | `hy-mt2-7b-q4tp.cmf` | bf16 checkpoint | 4.17 | 3.93 GB | 90.3⁴ | 98.6⁴ |
92
+ | Hy-MT2-30B-A3B | `hy-mt2-30b-a3b-q4tp.cmf` | bf16 checkpoint | 4.17 | 15.83 GB | 9.88 | 18.12 |
93
+
94
+ ¹ Twelve 512-token windows of wikitext-2 test, exact attention, the same
95
+ yardstick every card on this account uses. The bf16-class reference for
96
+ the 1.8B (`q8_2f`, not published) scores 24.77 — the 4-bit file
97
+ sits within 10% (27.13 vs 24.77) of it. These are translation models: their
98
+ perplexity on English encyclopedia prose is not what they were trained
99
+ for, which is why column ² exists.
100
+
101
+ ² Sixteen chat-formatted sentence pairs (RU/EN/ZH/DE/FR/ES/JA/IT/UK/PT,
102
+ 1 073 tokens) in the model's own prompt format — the number that tracks
103
+ what the file is for. The `q1t` file is Tencent's own 1.25-bit
104
+ quantization-aware checkpoint carried over **exactly** (see below); it
105
+ is a translation-only specialist and its general-text perplexity is not
106
+ meaningful, but its translations are — see the samples further down.
107
+
108
+ ³ Not a conversion defect: llama.cpp with the STQ1_0 kernel (PR #22836)
109
+ scores the very same GGUF at 1 793 on its own wikitext-2 protocol (first
110
+ twelve 512-token chunks, BOS-anchored, second halves scored), our
111
+ cold-window protocol at 4 460 — both say the 1.25-bit checkpoint no
112
+ longer models free English prose, while it translates correctly on every
113
+ backend. On the same Xeon that llama.cpp build decodes it at 2.2 tok/s;
114
+ cortiq's `q1t` kernels at 21.9.
115
+
116
+ ⁴ The 7B uses HunYuan's older tokenizer (vocab 128 167, GPT-4-style
117
+ splitting, `<|startoftext|>` … `<|extra_0|>` chat markup) and is sensitive
118
+ to running without its BOS in the middle of a corpus; its translations are
119
+ the best of the three dense files. Compare sizes on column ², not here.
120
+
121
+ Every q4tp file was quantized straight from the bf16 safetensors,
122
+ streamed shard by shard; nothing here is a re-quantization of a
123
+ lower-precision file. The tied `lm_head` of the dense models is not
124
+ written twice: the embedding serves both ends, as in the source.
125
+
126
+ ## What the runtime had to learn
127
+
128
+ Two things distinguish HunYuan from the Qwen/Llama block, and both ride
129
+ in the file header rather than in flags:
130
+
131
+ - **Dense 1.8B / 7B (`hunyuan_v1_dense`)** apply the per-head q/k RMSNorm
132
+ *after* RoPE (`query_layernorm` / `key_layernorm` on the rotated
133
+ vectors). That is not a reordering you can fold into the weights — the
134
+ rotation preserves a head's norm but not its elementwise norm weights —
135
+ so the engine carries an explicit `qk_norm_after_rope` flag through the
136
+ CPU, Metal and WGSL rope kernels. Their "dynamic" NTK-alpha RoPE
137
+ (`alpha = 1000`) is one rescaled base, `10000 · 1000^(128/126) =
138
+ 11 158 840`, applied at conversion; the full 262 144-token window is
139
+ native, nothing is rescaled per position.
140
+ - **30B-A3B (`hy_v3`)** is 48 layers with layer 0 dense (6 912 wide) and
141
+ 47 sparse layers of 128 experts, top-8 by *sigmoid* score plus a
142
+ selection bias (`expert_bias`, DeepSeek-V3 style: bias for the choice,
143
+ unbiased scores for the weights), renormalized and multiplied by
144
+ `router_scaling_factor = 2.826`, plus one always-on shared expert of
145
+ 768. The converter maps `mlp.router.gate` / `mlp.shared_mlp` onto the
146
+ canonical layout and the header carries the router constants; the wgpu
147
+ whole-token graph learned the two pieces it lacked — the routed scale
148
+ and a shared expert without a gate — so the 30B decodes on the card
149
+ as one graph (12 submits per token) instead of falling to the per-op
150
+ path (145 submits, 1.2 tok/s) the way every sigmoid-scaled MoE did
151
+ before 0.7.0.
152
+
153
+ ## The 1.25-bit file, exactly
154
+
155
+ Tencent ships `Hy-MT2-1.8B-1.25Bit-GGUF`: a quantization-aware ternary
156
+ checkpoint in llama.cpp's `STQ1_0` type (256-weight blocks, one f16
157
+ scale, and in every group of four lanes exactly one zero and three ±1 —
158
+ 5 bits per 4 weights). `cortiq import-gguf` decodes that block layout
159
+ natively and re-encodes each 32-group as CMF `q1t` (ternary, base-3
160
+ packed, f16 scale, empty outlier overlay): the reconstruction is
161
+ **bit-for-bit** the values the GGUF stores — no calibration, no second
162
+ quantizer. The cost of the general-purpose container is size: 2.25
163
+ bits/weight against 1.31, so the file is 695 MB against 462, with the
164
+ token table at `q8_2f` (Tencent stores it at 6.5 bits).
165
+
166
+ ```bash
167
+ cortiq import-gguf Hy-MT2-1.8B-1.25Bit.gguf --quant q1t \
168
+ --tokenizer-dir ./Hy-MT2-1.8B --output hy-mt2-1.8b-q1t.cmf
169
+ ```
170
+
171
+ `--tokenizer-dir` matters: HunYuan's pre-tokenizer splits digits in runs
172
+ of one to three and isolates CJK before the byte-level step, which a
173
+ tokenizer rebuilt from GGUF metadata cannot express. The flag embeds the
174
+ vendor `tokenizer.json` and chat template verbatim, so the ternary file
175
+ tokenizes exactly like the 4-bit ones.
176
+
177
+ ## Measured
178
+
179
+ Steady-state decode, single stream, `cortiq bench --core --tokens 128
180
+ --ignore-eos`, cortiq 0.7.0. The dense files are latency-bound on a
181
+ discrete card (a 1 GB model needs ~8 submits per token), so the CPU
182
+ matters as much as the GPU there; the MoE row is where the card counts.
183
+
184
+ | file | RTX PRO 4000 Blackwell 24 GB, Vulkan | Xeon E5-2690 v4 (14C), CPU | Apple M4 24 GB, Metal | M4, CPU (10 cores) |
185
+ |---|---:|---:|---:|---:|
186
+ | `hy-mt2-1.8b-q4tp.cmf` | 137.7 | 29.1 | 72.8 | 56.9 |
187
+ | `hy-mt2-1.8b-q1t.cmf` | 63.7 | 21.9 | 66.9 | 38.1 |
188
+ | `hy-mt2-7b-q4tp.cmf` | 77.2 | 9.3 | 23.1 | 16.7 |
189
+ | `hy-mt2-30b-a3b-q4tp.cmf` | 52.7 | 11.0 | —⁵ | — |
190
+
191
+ ### The MoE on any card (dynamic loading)
192
+
193
+ The 30B's whole-token graph is not all-or-nothing: as many leading layers
194
+ as the VRAM budget admits stay resident on the card, the host finishes the
195
+ rest — one boundary crossing per token, no expert paging. Measured on the
196
+ RTX PRO 4000 with `CMF_GPU_VRAM_MB` capped to what each card size would
197
+ auto-detect (the CPU alone: 10.5 tok/s on this 14-core Xeon):
198
+
199
+ | VRAM budget | 4 GB | 6 GB | 8 GB | 12 GB | 16 GB | 24 GB |
200
+ |---|---:|---:|---:|---:|---:|---:|
201
+ | layers on the card | 7/48 | 14/48 | 20/48 | 33/48 | 45/48 | **48/48** |
202
+ | decode, tok/s | 9.9 | 11.1 | 14.5 | 21.9 | 39.2 | **53.7** |
203
+
204
+ Perplexity and the greedy continuation do not move along the ladder — the
205
+ split changes where a layer runs, never what it computes. `CMF_GPU_VRAM_MB`
206
+ overrides the auto-detected budget when you want to cap it by hand.
207
+
208
+ ⁵ Not measured: the 15.8 GB file is at the edge of a 24 GB Mac (the 14.3 GB
209
+ Qwen3.8-27B decodes at 5.7 tok/s there) and on Metal a sigmoid-routed,
210
+ ungated-shared MoE layer still runs its experts on the CPU — the Metal
211
+ select kernel is the next port.
212
+
213
+ Prompt ingest (41-token prompt): on the RTX PRO 4000 the dense files take
214
+ 180 (1.8B) and 118 (7B) tok/s; the 30B ingests at **8 tok/s** on this card
215
+ — its batched prefill re-stages the expert buffers per 32-token chunk
216
+ instead of sharing the decode graph's resident copy, so a long source
217
+ paragraph costs seconds before the first token. Decode is unaffected;
218
+ sharing the buffers is the next item on the MoE list. On the M4: 469 tok/s
219
+ for the 1.8B q4tp, 218 for the ternary file, 122 for the 7B.
220
+
221
+ **First answer vs. the rest.** On a discrete card the weights are uploaded
222
+ when the whole-token graph is first built — 18.5 GB for the 30B, ~26 s on
223
+ this box — and `cortiq run` folds that into its printed "decode" figure
224
+ for a one-shot prompt, which then reads a few tok/s. The per-token cost
225
+ after the upload is the table above (`cortiq bench --core` measures it
226
+ after an untimed warm-up). `cortiq serve` pays the upload once per
227
+ process, so every request after the first decodes at the steady rate.
228
+
229
+ ## Prompting
230
+
231
+ The models have no system prompt. The chat template is embedded, so
232
+ `cortiq run` and `cortiq serve` wrap a plain user message correctly;
233
+ what goes into the message is Tencent's instruction, with the language
234
+ name written out in the prompt's language:
235
+
236
+ ```
237
+ Translate the following text into {target_lang}. Note that you should only output the translated result without any additional explanation:
238
+
239
+ {source_text}
240
+ ```
241
+
242
+ ```
243
+ 将以下文本翻译为{target_lang},注意只需要输出翻译后的结果,不要额外解释:
244
+
245
+ {source_text}
246
+ ```
247
+
248
+ Terminology, style, personalization, delimiter-preserving and
249
+ structured-data prompts are documented on the
250
+ [source card](https://huggingface.co/tencent/Hy-MT2-30B-A3B#hy-mt2-translation-task-instruction-examples-chinese-english-comparison).
251
+ Recommended sampling (Tencent): 1.8B / 7B — temperature 0.7, top-p 0.6,
252
+ top-k 20, repetition penalty 1.05; 30B-A3B — temperature 0.7, top-p 1.0,
253
+ no repetition penalty. `--greedy` is the deterministic choice for
254
+ evaluation.
255
+
256
+ Samples from the 1.8B files, greedy, identical on CPU and GPU:
257
+
258
+ | prompt | q4tp | q1t (1.25-bit) |
259
+ |---|---|---|
260
+ | RU→EN «Сегодня хорошая погода, и мы пойдём гулять в парк, а вечером посмотрим фильм.» | Today the weather is nice, and we're going to go for a walk in the park. In the evening, we'll watch a movie. | Today, the weather is nice. We'll go for a walk in the park, and in the evening, we'll watch a movie. |
261
+ | EN→DE "Please keep the {placeholder} and the number 12345 exactly as they are." | Bitte behalten Sie den {Placeholder} und die Zahl 12345 genau so, wie sie sind. | Bitte lassen Sie {placeholder} und die Zahl 12345 genau so, wie sie sind. |
262
+ | ZH→EN «今天天气真好,我们去公园散步吧。» | The weather is really nice today. Let's go for a walk in the park. | The weather is really nice today; let's go for a walk in the park. |
263
+
264
+ ## Server
265
+
266
+ ```bash
267
+ cortiq serve hy-mt2-7b-q4tp.cmf --port 8080
268
+ curl http://localhost:8080/v1/chat/completions -H 'Content-Type: application/json' \
269
+ -d '{"model":"hy-mt2-7b-q4tp","temperature":0,"messages":[{"role":"user","content":"Translate the following text into French. Note that you should only output the translated result without any additional explanation:\n\nThe meeting is at ten."}]}'
270
+ ```
271
+
272
+ `/v1/chat/completions`, `/v1/completions` and `/v1/models` are live;
273
+ `--ollama` adds an Ollama-compatible listener.
274
+
275
+ ## Reproduce
276
+
277
+ ```bash
278
+ cortiq convert --model tencent/Hy-MT2-1.8B --quant q4tp --output hy-mt2-1.8b-q4tp.cmf
279
+ cortiq convert --model tencent/Hy-MT2-7B --quant q4tp --output hy-mt2-7b-q4tp.cmf
280
+ cortiq convert --model tencent/Hy-MT2-30B-A3B --quant q4tp --output hy-mt2-30b-a3b-q4tp.cmf
281
+ ```
282
+
283
+ Streaming from the hub: peak disk is the output file. Every file was
284
+ size- and hash-verified after upload (`cortiq verify`).
285
+
286
+ Weights derive from Tencent's release and remain under its Apache-2.0
287
+ terms. The CMF container and the cortiq runtime are Apache-2.0 as well
288
+ (see the repository's LICENSE and PATENTS.md).