TokenAIzer commited on
Commit
2d29ae2
·
verified ·
1 Parent(s): 77ceb4c

Add model card

Browse files
Files changed (1) hide show
  1. README.md +218 -0
README.md ADDED
@@ -0,0 +1,218 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: XingChen-AGI/Xing4.0-29B-A4B
3
+ model_name: Xing4.0-29B-A4B-4bit-MLX
4
+ library_name: mlx
5
+ pipeline_tag: text-generation
6
+ license: apache-2.0
7
+ tags:
8
+ - mlx
9
+ - quantization
10
+ - apple-silicon
11
+ - moe
12
+ - mla
13
+ - hyper-connections
14
+ - xing
15
+ - telechat
16
+ - base_model:quantized:XingChen-AGI/Xing4.0-29B-A4B
17
+ - base_model_size:10B to 100B
18
+ ---
19
+
20
+ # Xing4.0-29B-A4B-4bit-MLX — 4-bit MLX quant of Xing4.0-29B-A4B
21
+
22
+ Unofficial Apple Silicon quantization of **[XingChen-AGI/Xing4.0-29B-A4B](https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B)**,
23
+ produced with `mlx_lm.convert` (MLX `affine` quantizer, group size 64) on an Apple M5 Max / 128 GB.
24
+
25
+ I am not affiliated with China Telecom AI. All upstream weights, benchmarks and license terms belong
26
+ to them, and the upstream Apache-2.0 license governs this repository too (see [License](#license)).
27
+
28
+ > **Format note:** these are **MLX safetensors**, not GGUF. They will not load in llama.cpp / Ollama /
29
+ > LM Studio, and they will not load in PyTorch/vLLM/SGLang either. Use an MLX runtime.
30
+
31
+ ## Before you download: you need `xing4_0` support in mlx-lm
32
+
33
+ Xing4.0 uses a new architecture (`model_type: xing4_0`) that **mlx-lm does not implement yet**.
34
+ Without it any MLX runtime stops with:
35
+
36
+ ```
37
+ ValueError: Model type xing4_0 not supported.
38
+ ```
39
+
40
+ An implementation exists and is verified against the upstream PyTorch code (see
41
+ [Provenance](#provenance)), but it is not merged into mlx-lm at the time of writing. Until it is,
42
+ these weights will not load anywhere. If you need them now, open an issue here and I will point you
43
+ at the model file.
44
+
45
+ ## Pick a variant
46
+
47
+ | | [4bit](https://huggingface.co/suzu89/Xing4.0-29B-A4B-4bit-MLX) | [6bit](https://huggingface.co/suzu89/Xing4.0-29B-A4B-6bit-MLX) | [8bit](https://huggingface.co/suzu89/Xing4.0-29B-A4B-8bit-MLX) |
48
+ |---|---|---|---|
49
+ | Weights on disk | 15.51 GiB (16.65 GB), 4 shards | 22.37 GiB (24.02 GB), 5 shards | 29.23 GiB (31.38 GB), 6 shards |
50
+ | Effective precision | 4.514 bits per weight | 6.512 bits per weight | 8.509 bits per weight |
51
+ | Peak RAM, short prompt | 15.6 GB | 22.4 GB | 29.3 GB |
52
+ | Measured generation | 21.2 tok/s | 21.0 tok/s | 21.1 tok/s |
53
+ | Choose when | prioritize memory headroom | balance size and weight precision | prioritize weight precision and have the memory |
54
+
55
+ All three drop the multi-token-prediction layer (see below). Throughput is nearly identical across
56
+ the three because only ~4B parameters are active per token; the difference shows up in memory, not
57
+ speed. Task-level accuracy after quantization has **not** been measured. Runtime memory also depends
58
+ on context length and KV cache.
59
+
60
+ ## What is inside (read from the shipped `config.json`)
61
+
62
+ | Field | Value |
63
+ |---|---|
64
+ | Architecture | `Xing4_0ForCausalLM` (`model_type: xing4_0`) |
65
+ | Parameters served | 29.51 B total, 4 B active per token |
66
+ | Layers / hidden | 40 layers, `hidden_size` 3584, dense FFN 9216 |
67
+ | Attention | MLA — `q_lora_rank` 768, `kv_lora_rank` 512, `qk_nope_head_dim` 128, `qk_rope_head_dim` 64, `v_head_dim` 128, 32 heads |
68
+ | MoE | 64 routed experts (`moe_intermediate_size` 1024) + 1 shared, 4 active per token, `noaux_tc` routing with sigmoid scoring, first 2 layers dense |
69
+ | Residual stream | mHC hyper-connections: `hc_mult` 4 parallel streams mixed by a Sinkhorn-normalized matrix (`hc_sinkhorn_iters` 20), two per layer |
70
+ | Position encoding | YaRN, `factor` 64 over `original_max_position_embeddings` 4096, interleaved RoPE |
71
+ | Context | `max_position_embeddings: 262144` |
72
+ | Vocab | 131,072 (tokenizer, `tokenization_xing4_0.py` and `chat_template.jinja` copied unchanged) |
73
+ | MTP | **dropped** — `num_nextn_predict_layers` normalized to 0 |
74
+
75
+ ## Quantization recipe
76
+
77
+ - `mlx_lm.convert -q --q-bits 4 --q-group-size 64`, mode `affine`, source BF16
78
+ - 476 modules quantized: attention projections, all 64 routed experts + shared expert per MoE layer, the dense FFNs of layers 0–1, `embed_tokens` and `lm_head`
79
+ - effective **4.514 bits per weight** (scales and biases included)
80
+ - **never quantized:** all RMSNorms, the MoE router (`mlp.gate`), the hyper-connection tables (`hc_fn`, `hc_base` in BF16), and `hc_scale` / `e_score_correction_bias` kept in FP32
81
+
82
+ ### About the dropped MTP layer
83
+
84
+ The upstream checkpoint carries a 41st layer (1.71 B parameters) implementing multi-token
85
+ prediction: `eh_proj`, `enorm`, `hnorm`, its own `embed_tokens`, a full attention + MoE block and a
86
+ `shared_head`. MLX has no speculative-decoding path for this architecture, so those tensors are
87
+ dropped and `num_nextn_predict_layers` is set to 0 to keep the shipped config self-consistent. If
88
+ you want MTP, use the upstream BF16 checkpoint with a runtime that supports it.
89
+
90
+ ## Requirements
91
+
92
+ - Apple Silicon, macOS 15+ (built and smoke-tested on M5 Max, 128 GB)
93
+ - `mlx-lm` **with `xing4_0` support** — see the warning above
94
+ - roughly 15.6 GB of free unified memory for a short prompt, more for long context
95
+
96
+ ## Usage
97
+
98
+ ```python
99
+ from mlx_lm import load, generate
100
+
101
+ # the custom tokenizer is loaded from the repo, so both flags are needed
102
+ model, tokenizer = load(
103
+ "suzu89/Xing4.0-29B-A4B-4bit-MLX",
104
+ tokenizer_config={"trust_remote_code": True},
105
+ trust_remote_code=True,
106
+ )
107
+
108
+ prompt = tokenizer.apply_chat_template(
109
+ [{"role": "user", "content": "Explain Sinkhorn normalization in one sentence."}],
110
+ add_generation_prompt=True,
111
+ )
112
+ print(generate(model, tokenizer, prompt=prompt, max_tokens=512, verbose=True))
113
+ ```
114
+
115
+ `mlx_lm.load` forwards `trust_remote_code` to the model but not to the tokenizer, which is why
116
+ `tokenizer_config` carries its own flag. Without it you get an unrelated-looking
117
+ `AttributeError: 'PreTrainedConfig' object has no attribute 'max_position_embeddings'`.
118
+
119
+ The chat template supports `enable_thinking` (on by default) and tool calls. The model tends to
120
+ write its reasoning trace in Chinese even for English prompts; that is upstream behaviour, not a
121
+ quantization artifact.
122
+
123
+ ### oMLX
124
+
125
+ oMLX 0.6.4 cannot run this architecture yet, and its oQ mixed-precision quantizer fails on it for
126
+ the same reason (its sensitivity pass cannot load the model). These are plain MLX quants, not oQ
127
+ builds.
128
+
129
+ ## Recommended sampling
130
+
131
+ Upstream recommends, and the shipped `generation_config.json` matches:
132
+
133
+ | Scenario | temperature | top_p | repetition_penalty |
134
+ |---|---|---|---|
135
+ | Complex reasoning / general | 1.0 | 0.95 | 1.05 |
136
+ | Coding / agent tasks | 0.8 | 0.95 | 1.05 |
137
+
138
+ Note that `mlx_lm.generate` does not apply a repetition penalty unless you pass a logits processor.
139
+
140
+ ## Provenance
141
+
142
+ The MLX implementation used to produce and load these weights was validated before quantizing:
143
+
144
+ | Check | Result |
145
+ |---|---|
146
+ | Hyper-connection vs the upstream PyTorch module, float32 | max relative error < 1e-5 |
147
+ | Full model vs `modeling_xing4_0.py`, small random-weight config | **max relative error 2.6e-07 on logits, 100% argmax agreement** |
148
+ | Upstream BF16 checkpoint, 58 GB | loads and generates coherent text |
149
+ | Each quant in this family | loads and generates coherent text |
150
+
151
+ Two upstream bugs found along the way, neither affecting these weights: the reference
152
+ `_init_weights` initializes `module.fn/base/scale` while the class defines `hc_fn/hc_base/hc_scale`
153
+ (random init from config fails, loading pretrained weights is unaffected), and `mlx_lm.load` does
154
+ not forward `trust_remote_code` to the tokenizer.
155
+
156
+ ## Benchmarks
157
+
158
+ I publish no numbers I have not measured myself. The table below is **upstream's**, measured on the
159
+ BF16 model, and is not a measurement of these quantized weights:
160
+
161
+ | Benchmark | Xing4.0-29B-A4B (BF16, upstream) | this quant |
162
+ |---|---|---|
163
+ | IFBench | 69.67 | not measured |
164
+ | AIME2026 | 90.00 | not measured |
165
+ | AA.LCR | 61.00 | not measured |
166
+ | Tau3-Bench | 64.63 | not measured |
167
+ | Claw-Eval | 76.55 | not measured |
168
+ | SWE-bench Verified | 75.00 | not measured |
169
+ | Terminal-Bench 2.1 | 57.50 | not measured |
170
+ | SWE-bench Multilingual | 66.00 | not measured |
171
+ | DeepresearchBII | 60.80 | not measured |
172
+
173
+ Measurements and issue reports ("quant X broke task Y") are welcome and will be merged into this
174
+ table.
175
+
176
+ ## Known caveats
177
+
178
+ - Quantization is lossy. If you see a regression, compare against a higher-precision variant and the
179
+ BF16 source before filing a bug.
180
+ - The hyper-connection mixing runs in the residual path of every layer and is kept in BF16/FP32 here;
181
+ its sensitivity to weight quantization elsewhere in the model has not been studied.
182
+ - 262k context is the architecture's limit, not a promise: keep the KV cache inside your memory
183
+ budget or the machine swaps.
184
+ - No MTP head, so no self-speculative decoding.
185
+ - Agentic and long-context behaviour at 4-bit is untested.
186
+
187
+ ## License
188
+
189
+ Distributed under the **Apache License 2.0**, inherited from
190
+ [XingChen-AGI/Xing4.0-29B-A4B](https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B). See `LICENSE-NOTICE.md` in this repository.
191
+
192
+ ## Citation
193
+
194
+ ```bibtex
195
+ @misc{xing4-29b-a4b-mlx-4bit,
196
+ title = {Xing4.0-29B-A4B-4bit-MLX: MLX 4-bit quantization of Xing4.0-29B-A4B},
197
+ author = {suzu89},
198
+ year = {2026},
199
+ howpublished = {\url{https://huggingface.co/suzu89/Xing4.0-29B-A4B-4bit-MLX}},
200
+ note = {Unofficial quantization of XingChen-AGI/Xing4.0-29B-A4B}
201
+ }
202
+
203
+ @misc{liu2025trainingreporttelechat3moe,
204
+ title = {Training Report of TeleChat3-MoE},
205
+ author = {Xinzhang Liu and others},
206
+ year = {2025},
207
+ eprint = {2512.24157},
208
+ archivePrefix = {arXiv},
209
+ primaryClass = {cs.CL},
210
+ url = {https://arxiv.org/abs/2512.24157}
211
+ }
212
+ ```
213
+
214
+ ## Acknowledgements
215
+
216
+ - **China Telecom AI (XingChen-AGI)** for Xing4.0-29B-A4B and the mHC architecture.
217
+ - **Apple MLX team** for `mlx` and `mlx-lm`, whose DeepSeek-V3 implementation this architecture
218
+ builds on directly.