npario Youssofal commited on
Commit
8e8a717
·
0 Parent(s):

Duplicate from Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality

Browse files

Co-authored-by: Youssof Altoukhi <Youssofal@users.noreply.huggingface.co>

.gitattributes ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ *.7z filter=lfs diff=lfs merge=lfs -text
2
+ *.arrow filter=lfs diff=lfs merge=lfs -text
3
+ *.bin filter=lfs diff=lfs merge=lfs -text
4
+ *.bz2 filter=lfs diff=lfs merge=lfs -text
5
+ *.ckpt filter=lfs diff=lfs merge=lfs -text
6
+ *.ftz filter=lfs diff=lfs merge=lfs -text
7
+ *.gz filter=lfs diff=lfs merge=lfs -text
8
+ *.h5 filter=lfs diff=lfs merge=lfs -text
9
+ *.joblib filter=lfs diff=lfs merge=lfs -text
10
+ *.lfs.* filter=lfs diff=lfs merge=lfs -text
11
+ *.mlmodel filter=lfs diff=lfs merge=lfs -text
12
+ *.model filter=lfs diff=lfs merge=lfs -text
13
+ *.msgpack filter=lfs diff=lfs merge=lfs -text
14
+ *.npy filter=lfs diff=lfs merge=lfs -text
15
+ *.npz filter=lfs diff=lfs merge=lfs -text
16
+ *.onnx filter=lfs diff=lfs merge=lfs -text
17
+ *.ot filter=lfs diff=lfs merge=lfs -text
18
+ *.parquet filter=lfs diff=lfs merge=lfs -text
19
+ *.pb filter=lfs diff=lfs merge=lfs -text
20
+ *.pickle filter=lfs diff=lfs merge=lfs -text
21
+ *.pkl filter=lfs diff=lfs merge=lfs -text
22
+ *.pt filter=lfs diff=lfs merge=lfs -text
23
+ *.pth filter=lfs diff=lfs merge=lfs -text
24
+ *.rar filter=lfs diff=lfs merge=lfs -text
25
+ *.safetensors filter=lfs diff=lfs merge=lfs -text
26
+ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
27
+ *.tar.* filter=lfs diff=lfs merge=lfs -text
28
+ *.tar filter=lfs diff=lfs merge=lfs -text
29
+ *.tflite filter=lfs diff=lfs merge=lfs -text
30
+ *.tgz filter=lfs diff=lfs merge=lfs -text
31
+ *.wasm filter=lfs diff=lfs merge=lfs -text
32
+ *.xz filter=lfs diff=lfs merge=lfs -text
33
+ *.zip filter=lfs diff=lfs merge=lfs -text
34
+ *.zst filter=lfs diff=lfs merge=lfs -text
35
+ *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,123 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: mtplx
4
+ pipeline_tag: text-generation
5
+ base_model: Qwen/Qwen3.8-27B
6
+ base_model_relation: quantized
7
+ tags:
8
+ - mlx
9
+ - apple-silicon
10
+ - macos
11
+ - speculative-decoding
12
+ - multi-token-prediction
13
+ - qwen
14
+ - qwen3.8
15
+ - mtp
16
+ - mtplx
17
+ - local-ai
18
+ - coding
19
+ - qwen3-8
20
+ - qwen-3.8
21
+ - local-llm
22
+ - llm
23
+ - m5
24
+ - m5-max
25
+ - m4
26
+ - m3
27
+ - macbook-pro
28
+ - mac-studio
29
+ - opencode
30
+ - claude-code
31
+ - 27b
32
+ - qwen3.8-27b
33
+ - qwen3-8-27b
34
+ ---
35
+
36
+ **[MTPLX](https://mtplx.com): the fastest way to run Qwen 3.8 on a Mac. Native multi-token-prediction speculative decoding on Apple Silicon, two to three times the speed of plain decoding, exact at any temperature.**
37
+
38
+ # Qwen 3.8 27B Optimized Quality
39
+
40
+ **8-bit dynamic quant. Good coding speeds and perfect quality.**
41
+
42
+ The closest of the three MTPLX Qwen 3.8 builds to the original bf16 model.
43
+ Every weight matrix at 8-bit, sensitive parts at 16-bit, native
44
+ multi-token-prediction head kept, so [MTPLX](https://mtplx.com) still drafts
45
+ ahead and verifies in one pass. Pick this when you want the answer the full
46
+ model would give and still want it fast.
47
+
48
+ ## Measured on MTPLX 2.11.3 (16 September 2026)
49
+
50
+ This is a Qwen3.8-27B MLX pack for MTPLX, the fastest way to run Qwen 3.8 27B on a Mac. All numbers on a MacBook Pro M5 Max, fans verified at maximum, sampled at the model's own settings. Conditions and sources: [mtplx.com/benchmarks](https://mtplx.com/benchmarks/).
51
+
52
+ | Run | tok/s |
53
+ |---|---|
54
+ | Optimized Speed rewriting a file it just wrote, stock settings (MTPLX 2.10.0) | 87.6 |
55
+ | Bare Speed on a fresh coding task at official Qwen 3.8 sampling (MTPLX 2.7.0) | 65.2 |
56
+ | Optimized Speed on a 3k-token chat answer (MTPLX 2.10.0) | 64.3 |
57
+ | The 27B record on a fresh generation: Qwen 3.6 27B Optimized Speed, 192-token bench, 2 July 2026, raw logs published | 81.74 |
58
+
59
+ Pack quality against the bf16 Qwen3.8-27B checkpoint, teacher-forced over 2,389 positions of code, prose, JSON and a multilingual notice (16 September 2026): Optimized Speed (4-bit dynamic) 96.0 percent top-1 agreement and KL 0.012; Optimized Quality (8-bit dynamic) 99.3 percent and KL 0.0005. Exactness on this release: a thousand four-token draws from the fast path match a thousand from the plain path within the plain path's own noise, at temperature 1, top-p 0.95, top-k 20. Details: [MTPLX 2.11.3 release notes](https://mtplx.com/releases/2.11.3/).
60
+
61
+ Runs on Apple Silicon Macs with 32 GB of unified memory or more (36 GB for Optimized Quality): MacBook Pro, MacBook Air, Mac mini and Mac Studio on M1 to M5. Guide: [Run Qwen 3.8 27B on a Mac](https://mtplx.com/models/qwen3.8-27b/). Comparison pages: [MTPLX vs mlx-serve](https://mtplx.com/compare/mtplx-vs-mlx-serve/), [MTPLX vs oMLX](https://mtplx.com/compare/mtplx-vs-omlx/), [MTPLX vs LM Studio](https://mtplx.com/compare/mtplx-vs-lm-studio/), [MTPLX vs Ollama](https://mtplx.com/compare/mtplx-vs-ollama/).
62
+
63
+ ## Speeds
64
+
65
+ Measured on an M5 Max, fans verified at max, single stream, generation running
66
+ to the model's own stop, official Qwen 3.8 sampling (temperature 1.0, top-p
67
+ 0.95, top-k 20).
68
+
69
+ | Run | tok/s |
70
+ |---|---|
71
+ | Coding task, medium reasoning, inside the MTPLX Mac app | 48.3 |
72
+ | Long reasoning at xhigh, 34k and 46k token answers | 33.2 and 33.1 |
73
+
74
+ Same night, same task, the 4-bit builds: Qwen 3.6 27B Optimized Speed V2 59.9
75
+ to 60.1 tok/s, Qwen 3.8 Optimized Speed 58.7, Bare Speed 65.2. This is the
76
+ quality pick, not the speed pick, and it is still well past 40 tok/s while
77
+ running the full-precision distribution.
78
+
79
+ Draft acceptance on the coding task by depth: 0.96, 0.88, 0.79. Verify cost
80
+ 63.5 ms per round. Depth 3 was measured at +19.9% over depth 2 on the long
81
+ reasoning task.
82
+
83
+ ## How it is built
84
+
85
+ - Every weight matrix at 8-bit with 64-weight groups.
86
+ - The GDN convolution kernels and recurrent state parameters, every norm, and
87
+ the whole MTP head stay 16-bit.
88
+ - KL divergence to the original bf16 model on our coding battery: 0.00105.
89
+ That is 21x closer than Optimized Speed and 36x closer than Bare Speed. In
90
+ practice you will not tell the outputs apart from the bf16 model.
91
+
92
+ | | |
93
+ |---|---|
94
+ | Download | 29.4 GB |
95
+ | Peak unified memory (measured, this artifact) | 32.7 GB |
96
+ | Context window | 262,144 tokens |
97
+ | MTP depth | 3 |
98
+ | Sampling | temperature 1.0, top-p 0.95, top-k 20 (the official Qwen 3.8 contract) |
99
+
100
+ The tuned depth and draft settings ship inside `mtplx_runtime.json`. MTPLX
101
+ reads them on load. Speculation is exact: drafts are accepted with the
102
+ probability-ratio rule plus residual resampling, so the output follows the
103
+ model's own distribution at any temperature. Reasoning effort levels (xhigh,
104
+ medium, low) work, and preserved thinking flows through the MTP path.
105
+
106
+ ## Use it
107
+
108
+ You want 36 GB of unified memory or more for this one. Mac app: download at
109
+ [mtplx.com](https://mtplx.com), pick "Qwen 3.8 27B Optimized Quality".
110
+
111
+ Command line:
112
+
113
+ ```bash
114
+ pip install mtplx
115
+ mtplx serve --model Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality
116
+ ```
117
+
118
+ Siblings: [Optimized Speed](https://huggingface.co/Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed)
119
+ (recommended for coding) and
120
+ [Bare Speed](https://huggingface.co/Youssofal/Qwen3.8-27B-MTPLX-Bare-Speed)
121
+ (quickest burst chat speeds). On an M1 or M2 Mac use the
122
+ [FP16 build](https://huggingface.co/Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality-FP16)
123
+ of this model.
chat_template.jinja ADDED
@@ -0,0 +1,170 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- set image_count = namespace(value=0) %}
2
+ {%- set video_count = namespace(value=0) %}
3
+ {%- macro render_content(content, do_vision_count, is_system_content=false) %}
4
+ {%- if content is string %}
5
+ {{- content }}
6
+ {%- elif content is iterable and content is not mapping %}
7
+ {%- for item in content %}
8
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
9
+ {%- if is_system_content %}
10
+ {{- raise_exception('System message cannot contain images.') }}
11
+ {%- endif %}
12
+ {%- if do_vision_count %}
13
+ {%- set image_count.value = image_count.value + 1 %}
14
+ {%- endif %}
15
+ {%- if add_vision_id %}
16
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
17
+ {%- endif %}
18
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
19
+ {%- elif 'video' in item or item.type == 'video' %}
20
+ {%- if is_system_content %}
21
+ {{- raise_exception('System message cannot contain videos.') }}
22
+ {%- endif %}
23
+ {%- if do_vision_count %}
24
+ {%- set video_count.value = video_count.value + 1 %}
25
+ {%- endif %}
26
+ {%- if add_vision_id %}
27
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
28
+ {%- endif %}
29
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
30
+ {%- elif 'text' in item %}
31
+ {{- item.text }}
32
+ {%- else %}
33
+ {{- raise_exception('Unexpected item type in content.') }}
34
+ {%- endif %}
35
+ {%- endfor %}
36
+ {%- elif content is none or content is undefined %}
37
+ {{- '' }}
38
+ {%- else %}
39
+ {{- raise_exception('Unexpected content type.') }}
40
+ {%- endif %}
41
+ {%- endmacro %}
42
+ {%- if not messages %}
43
+ {{- raise_exception('No messages provided.') }}
44
+ {%- endif %}
45
+ {%- set reasoning_instructions = '' %}
46
+ {%- if enable_thinking is undefined or enable_thinking is true %}
47
+ {%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}
48
+ {%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %}
49
+ {{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }}
50
+ {%- endif %}
51
+ {%- if resolved_reasoning_effort == 'xhigh' %}
52
+ {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}
53
+ {%- elif resolved_reasoning_effort == 'low' %}
54
+ {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}
55
+ {%- endif %}
56
+ {%- endif %}
57
+ {%- if tools and tools is iterable and tools is not mapping %}
58
+ {{- '<|im_start|>system\n' }}
59
+ {%- if reasoning_instructions %}
60
+ {{- reasoning_instructions + '\n\n' }}
61
+ {%- endif %}
62
+ {{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
63
+ {%- for tool in tools %}
64
+ {{- "\n" }}
65
+ {{- tool | tojson }}
66
+ {%- endfor %}
67
+ {{- "\n</tools>" }}
68
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
69
+ {%- if messages[0].role == 'system' %}
70
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
71
+ {%- if content %}
72
+ {{- '\n\n' + content }}
73
+ {%- endif %}
74
+ {%- endif %}
75
+ {{- '<|im_end|>\n' }}
76
+ {%- else %}
77
+ {%- if messages[0].role == 'system' %}
78
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
79
+ {%- if content %}
80
+ {{- '<|im_start|>system\n' + (reasoning_instructions + '\n\n' if reasoning_instructions else '') + content + '<|im_end|>\n' }}
81
+ {%- elif reasoning_instructions %}
82
+ {{- '<|im_start|>system\n' + reasoning_instructions + '<|im_end|>\n' }}
83
+ {%- endif %}
84
+ {%- elif reasoning_instructions %}
85
+ {{- '<|im_start|>system\n' + reasoning_instructions + '<|im_end|>\n' }}
86
+ {%- endif %}
87
+ {%- endif %}
88
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
89
+ {%- for message in messages[::-1] %}
90
+ {%- set index = (messages|length - 1) - loop.index0 %}
91
+ {%- if ns.multi_step_tool and message.role == "user" %}
92
+ {%- set content = render_content(message.content, false)|trim %}
93
+ {%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
94
+ {%- set ns.multi_step_tool = false %}
95
+ {%- set ns.last_query_index = index %}
96
+ {%- endif %}
97
+ {%- endif %}
98
+ {%- endfor %}
99
+ {%- if ns.multi_step_tool %}
100
+ {{- raise_exception('No user query found in messages.') }}
101
+ {%- endif %}
102
+ {%- for message in messages %}
103
+ {%- set content = render_content(message.content, true)|trim %}
104
+ {%- if message.role == "system" %}
105
+ {%- if not loop.first %}
106
+ {{- raise_exception('System message must be at the beginning.') }}
107
+ {%- endif %}
108
+ {%- elif message.role == "user" %}
109
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
110
+ {%- elif message.role == "assistant" %}
111
+ {%- set reasoning_content = '' %}
112
+ {%- if message.reasoning_content is string %}
113
+ {%- set reasoning_content = message.reasoning_content %}
114
+ {%- endif %}
115
+ {%- set reasoning_content = reasoning_content|trim %}
116
+ {%- if preserve_thinking is undefined or preserve_thinking is true or loop.index0 > ns.last_query_index %}
117
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
118
+ {%- else %}
119
+ {{- '<|im_start|>' + message.role + '\n' + content }}
120
+ {%- endif %}
121
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
122
+ {%- for tool_call in message.tool_calls %}
123
+ {%- if tool_call.function is defined %}
124
+ {%- set tool_call = tool_call.function %}
125
+ {%- endif %}
126
+ {%- if loop.first %}
127
+ {%- if content|trim %}
128
+ {{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
129
+ {%- else %}
130
+ {{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
131
+ {%- endif %}
132
+ {%- else %}
133
+ {{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
134
+ {%- endif %}
135
+ {%- if tool_call.arguments is defined and tool_call.arguments != '' %}
136
+ {%- for args_name, args_value in tool_call.arguments|items %}
137
+ {{- '<parameter=' + args_name + '>\n' }}
138
+ {%- set args_value = args_value | string if args_value is string else args_value | tojson | safe %}
139
+ {{- args_value }}
140
+ {{- '\n</parameter>\n' }}
141
+ {%- endfor %}
142
+ {%- endif %}
143
+ {{- '</function>\n</tool_call>' }}
144
+ {%- endfor %}
145
+ {%- endif %}
146
+ {{- '<|im_end|>\n' }}
147
+ {%- elif message.role == "tool" %}
148
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
149
+ {{- '<|im_start|>user' }}
150
+ {%- endif %}
151
+ {{- '\n<tool_response>\n' }}
152
+ {{- content }}
153
+ {{- '\n</tool_response>' }}
154
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
155
+ {{- '<|im_end|>\n' }}
156
+ {%- elif loop.last %}
157
+ {{- '<|im_end|>\n' }}
158
+ {%- endif %}
159
+ {%- else %}
160
+ {{- raise_exception('Unexpected message role.') }}
161
+ {%- endif %}
162
+ {%- endfor %}
163
+ {%- if add_generation_prompt %}
164
+ {{- '<|im_start|>assistant\n' }}
165
+ {%- if enable_thinking is defined and enable_thinking is false %}
166
+ {{- '<think>\n\n</think>\n\n' }}
167
+ {%- else %}
168
+ {{- '<think>\n' }}
169
+ {%- endif %}
170
+ {%- endif %}
config.json ADDED
@@ -0,0 +1,184 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Qwen3_5ForConditionalGeneration"
4
+ ],
5
+ "eos_token_id": [
6
+ 248046,
7
+ 248044
8
+ ],
9
+ "image_token_id": 248056,
10
+ "language_model_only": false,
11
+ "mlx_lm_extra_tensors": {
12
+ "mtp_file": "mtp.safetensors"
13
+ },
14
+ "model_type": "qwen3_5",
15
+ "mtplx_mtp_contract": {
16
+ "base_hidden_variant": "post_norm",
17
+ "concat_order": "embedding_hidden",
18
+ "hidden_variant": "post_norm",
19
+ "mtp_position_mode": "local",
20
+ "mtp_quant_group_size": 64,
21
+ "mtp_quant_mode": "affine"
22
+ },
23
+ "mtplx_mtp_payload_audit": {
24
+ "mtp_file": "/Users/youssof/.mtplx/models/Qwen3.8-27B-MTPLX-Optimized-Quality/mtp.safetensors",
25
+ "nonzero_payload_tensor_count": 8,
26
+ "passed": true,
27
+ "payload_tensor_count": 8,
28
+ "problems": [],
29
+ "scale_tensor_count": 0,
30
+ "tensor_count": 15,
31
+ "zero_payload_sample": [],
32
+ "zero_scale_sample": []
33
+ },
34
+ "quantization": {
35
+ "bits": 8,
36
+ "group_size": 64,
37
+ "mode": "affine"
38
+ },
39
+ "quantization_config": {
40
+ "bits": 8,
41
+ "group_size": 64,
42
+ "mode": "affine"
43
+ },
44
+ "text_config": {
45
+ "attention_bias": false,
46
+ "attention_dropout": 0.0,
47
+ "attn_output_gate": true,
48
+ "bos_token_id": 248044,
49
+ "dtype": "bfloat16",
50
+ "eos_token_id": 248044,
51
+ "full_attention_interval": 4,
52
+ "head_dim": 256,
53
+ "hidden_act": "silu",
54
+ "hidden_size": 5120,
55
+ "initializer_range": 0.02,
56
+ "intermediate_size": 17408,
57
+ "layer_types": [
58
+ "linear_attention",
59
+ "linear_attention",
60
+ "linear_attention",
61
+ "full_attention",
62
+ "linear_attention",
63
+ "linear_attention",
64
+ "linear_attention",
65
+ "full_attention",
66
+ "linear_attention",
67
+ "linear_attention",
68
+ "linear_attention",
69
+ "full_attention",
70
+ "linear_attention",
71
+ "linear_attention",
72
+ "linear_attention",
73
+ "full_attention",
74
+ "linear_attention",
75
+ "linear_attention",
76
+ "linear_attention",
77
+ "full_attention",
78
+ "linear_attention",
79
+ "linear_attention",
80
+ "linear_attention",
81
+ "full_attention",
82
+ "linear_attention",
83
+ "linear_attention",
84
+ "linear_attention",
85
+ "full_attention",
86
+ "linear_attention",
87
+ "linear_attention",
88
+ "linear_attention",
89
+ "full_attention",
90
+ "linear_attention",
91
+ "linear_attention",
92
+ "linear_attention",
93
+ "full_attention",
94
+ "linear_attention",
95
+ "linear_attention",
96
+ "linear_attention",
97
+ "full_attention",
98
+ "linear_attention",
99
+ "linear_attention",
100
+ "linear_attention",
101
+ "full_attention",
102
+ "linear_attention",
103
+ "linear_attention",
104
+ "linear_attention",
105
+ "full_attention",
106
+ "linear_attention",
107
+ "linear_attention",
108
+ "linear_attention",
109
+ "full_attention",
110
+ "linear_attention",
111
+ "linear_attention",
112
+ "linear_attention",
113
+ "full_attention",
114
+ "linear_attention",
115
+ "linear_attention",
116
+ "linear_attention",
117
+ "full_attention",
118
+ "linear_attention",
119
+ "linear_attention",
120
+ "linear_attention",
121
+ "full_attention"
122
+ ],
123
+ "linear_conv_kernel_dim": 4,
124
+ "linear_key_head_dim": 128,
125
+ "linear_num_key_heads": 16,
126
+ "linear_num_value_heads": 48,
127
+ "linear_value_head_dim": 128,
128
+ "mamba_ssm_dtype": "float32",
129
+ "max_position_embeddings": 262144,
130
+ "model_type": "qwen3_5_text",
131
+ "mtp_num_hidden_layers": 1,
132
+ "mtp_use_dedicated_embeddings": false,
133
+ "num_attention_heads": 24,
134
+ "num_hidden_layers": 64,
135
+ "num_key_value_heads": 4,
136
+ "output_gate_type": "swish",
137
+ "pad_token_id": null,
138
+ "partial_rotary_factor": 0.25,
139
+ "rms_norm_eps": 1e-06,
140
+ "rope_parameters": {
141
+ "mrope_interleaved": true,
142
+ "mrope_section": [
143
+ 11,
144
+ 11,
145
+ 10
146
+ ],
147
+ "partial_rotary_factor": 0.25,
148
+ "rope_theta": 10000000,
149
+ "type": "default"
150
+ },
151
+ "tie_word_embeddings": false,
152
+ "use_cache": true,
153
+ "vocab_size": 248320
154
+ },
155
+ "tie_word_embeddings": false,
156
+ "transformers_version": "5.8.0.dev0",
157
+ "video_token_id": 248057,
158
+ "vision_config": {
159
+ "deepstack_visual_indexes": [],
160
+ "depth": 27,
161
+ "hidden_act": "gelu_pytorch_tanh",
162
+ "hidden_size": 1152,
163
+ "in_channels": 3,
164
+ "initializer_range": 0.02,
165
+ "intermediate_size": 4304,
166
+ "model_type": "qwen3_5",
167
+ "num_heads": 16,
168
+ "num_position_embeddings": 2304,
169
+ "out_hidden_size": 5120,
170
+ "patch_size": 16,
171
+ "spatial_merge_size": 2,
172
+ "temporal_patch_size": 2
173
+ },
174
+ "vision_end_token_id": 248054,
175
+ "vision_start_token_id": 248053,
176
+ "mtplx_mtp_quantization": {
177
+ "bits": 8,
178
+ "group_size": 64,
179
+ "mode": "affine",
180
+ "policy": "all",
181
+ "prequantized": true,
182
+ "description": "All 8 MTP draft-head matrices (fc + attention q/k/v/o + MLP gate/up/down) packed MLX INT8/g64 affine from the released sidecar; head norms keep the pack's float dtype. Verified flat-or-better acceptance vs the unquantized head before publishing."
183
+ }
184
+ }
generation_config.json ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 248044,
3
+ "do_sample": true,
4
+ "eos_token_id": [
5
+ 248046,
6
+ 248044
7
+ ],
8
+ "pad_token_id": 248044,
9
+ "temperature": 1.0,
10
+ "top_k": 20,
11
+ "top_p": 0.95
12
+ }
model-00001-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a11adc73146b142e77349999f67d5d13c031ff2bbafed0837388c1b6e48d4836
3
+ size 5305469313
model-00002-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bd95f804980c5e79e212abe6e674020a30f24674780a92cb101dd9971484c6d0
3
+ size 5354102628
model-00003-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:91e3f7cc9fadd3599ebce87a9f8e8a4ab6d05f595a79da94868fa96c2d868b8e
3
+ size 5337392986
model-00004-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ae6fee0e32edbf973e614ac040ad0f853df23f8c33bf97cadb4c168ef9a887a2
3
+ size 5292847128
model-00005-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c1735ac9c3d6cba8127b5c00a28576f0825d874a7f09911dc7bad5027a8747eb
3
+ size 5354102635
model-00006-of-00006.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2452a44eddf64d0cd964245b37f47f3b406c9c397b0eb5fd605bfb9c2f8825f9
3
+ size 1935805825
model-vision.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:964bef26c740bdb6fe464b4c7c48840d4c952ea597c181c41cb131c65ea3c5d5
3
+ size 921497225
model.safetensors.index.json ADDED
The diff for this file is too large to render. See raw diff
 
mtp.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1e350715e0509932e337eb60c074a3e47be1a9a04ea0fcb6c29b193b67c470c5
3
+ size 451270903
mtplx_runtime.json ADDED
@@ -0,0 +1,179 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "arch_id": "qwen3-next-mtp",
3
+ "artifact_role": "forge-local",
4
+ "base_trunk": "Qwen/Qwen3.8-27B",
5
+ "exactness_baseline": {},
6
+ "forge_provenance": {
7
+ "depth_default_note": "Depth pin REMOVED 2026-08-14 after the gated live ABBA: D2 39.5 vs D3 40.6 tok/s matched-window (tie; D3 fewer verify rounds 697 vs 883). The earlier forge-verify \"D3 18.8 collapse\" was order/JIT-confounded (mistakes/single-forge-verify-tune-rows...). Family ceiling D3 applies.",
8
+ "forge_inputs": {
9
+ "mtp_source_path": "<redacted>/Qwen--Qwen3.8-27B",
10
+ "trunk_path": "<redacted>/Qwen3.8-27B-MTPLX-Optimized-Quality"
11
+ },
12
+ "forge_recipe": {
13
+ "body_bits": 8,
14
+ "body_group_size": 64,
15
+ "body_mode": "affine",
16
+ "mtp_policy": "keep_bf16"
17
+ },
18
+ "forged_at": "2026-08-14T11:23:31-07:00",
19
+ "forged_locally": true,
20
+ "head_quantization": {
21
+ "bits": 8,
22
+ "group_size": 64,
23
+ "mode": "affine",
24
+ "note": "Structural head quantization of the released sidecar; no calibration, no training. Trunk weights unchanged.",
25
+ "policy": "all",
26
+ "quantized_at": "2026-08-20T04:49:56Z",
27
+ "quantized_sidecar_bytes": 451270903,
28
+ "source_sidecar_bytes": 849400403,
29
+ "tool": "scripts/build_qwen38_q4head_sidecar.py"
30
+ },
31
+ "mtp_contract": {
32
+ "base_hidden_variant": "post_norm",
33
+ "concat_order": "embedding_hidden",
34
+ "hidden_variant": "post_norm",
35
+ "mtp_position_mode": "local",
36
+ "mtp_quant_group_size": 64,
37
+ "mtp_quant_mode": "affine"
38
+ },
39
+ "mtplx_version": "2.6.0",
40
+ "published_to_hf": null,
41
+ "recommended_profile_note": "Restamped sustained->turbo 2026-08-14: the sustained recommendation came from the same confounded verify rows; every gated arm served turbo, medium Flappy 59.8 tok/s blended quiet-window.",
42
+ "source_format": "bf16_native",
43
+ "source_repo": "Qwen/Qwen3.8-27B",
44
+ "source_sha": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0"
45
+ },
46
+ "mtp_contract": {
47
+ "base_hidden_variant": "post_norm",
48
+ "concat_order": "embedding_hidden",
49
+ "hidden_variant": "post_norm",
50
+ "mtp_position_mode": "local",
51
+ "mtp_quant_group_size": 64,
52
+ "mtp_quant_mode": "affine"
53
+ },
54
+ "mtp_depth_default": 3,
55
+ "mtp_depth_max": 3,
56
+ "mtp_sidecar": "int8-g64-prequantized",
57
+ "mtplx_version": "2.9.0",
58
+ "recommended_draft_sampler": {
59
+ "temperature": 1.0,
60
+ "top_k": 20,
61
+ "top_p": 0.95
62
+ },
63
+ "recommended_profile": "turbo",
64
+ "release_validation": {
65
+ "draft_sampler_transfer_note": {
66
+ "basis": "dropday uncapped ABBA campaign (clean-receipts.jsonl, temp 1.0 xhigh/medium, die-gated, max fans verified)",
67
+ "note": "draft 1.0 adopted by transfer: OS uncapped ABBA (+13%) + Bare scaffold strict A/B (keep_official_draft_sampler_1_0); same MTP head architecture and temp-1.0 target contract. OQ depth arms measured at draft 0.6; app-receipt validates live at ship config.",
68
+ "stamped_at": "2026-08-14T21:45:00-07:00"
69
+ },
70
+ "dropday_depth_abba": {
71
+ "basis": "dropday uncapped ABBA campaign (clean-receipts.jsonl, temp 1.0 xhigh/medium, die-gated, max fans verified)",
72
+ "decode_tps_blended_d2": [
73
+ 27.3,
74
+ 28.0
75
+ ],
76
+ "decode_tps_blended_d3": [
77
+ 33.2,
78
+ 33.1
79
+ ],
80
+ "stamped_at": "2026-08-14T21:45:00-07:00",
81
+ "verdict": "depth_3 (+19.9% blended)",
82
+ "workload": "simple flappy xhigh uncapped (33k-46k tokens, EOS stop)"
83
+ }
84
+ },
85
+ "sampler": {
86
+ "temperature": 0.6,
87
+ "top_k": 20,
88
+ "top_p": 0.95
89
+ },
90
+ "speed_evidence": {
91
+ "acceptance_by_depth": [
92
+ 0.9433962264150944,
93
+ 0.8814016172506739,
94
+ 0.8059299191374663
95
+ ],
96
+ "acceptance_collapsed": [],
97
+ "artifact_fingerprint": "sha256:5a7a0d400ebff91dc544f7d93454d6d48d2833b4f0f7ad832d00e1d9513fd750",
98
+ "depth": 3,
99
+ "failure_reasons": [],
100
+ "forge_verify_rows": [
101
+ {
102
+ "acceptance_by_position": [],
103
+ "depth": 0,
104
+ "finish_reasons": {
105
+ "stop": 1
106
+ },
107
+ "hit_token_budget": false,
108
+ "hit_token_budget_count": 0,
109
+ "multiplier_vs_ar": 1.0,
110
+ "quality_passed": true,
111
+ "tok_s": 13.054796369549772,
112
+ "verify_time_s": 154.4356367919827
113
+ },
114
+ {
115
+ "acceptance_by_position": [
116
+ 0.9570405727923628
117
+ ],
118
+ "depth": 1,
119
+ "finish_reasons": {
120
+ "stop": 1
121
+ },
122
+ "hit_token_budget": false,
123
+ "hit_token_budget_count": 0,
124
+ "multiplier_vs_ar": 2.0262953935160954,
125
+ "quality_passed": true,
126
+ "tok_s": 26.45287374690935,
127
+ "verify_time_s": 35.17482435004786
128
+ },
129
+ {
130
+ "acceptance_by_position": [
131
+ 0.9776119402985075,
132
+ 0.9129353233830846
133
+ ],
134
+ "depth": 2,
135
+ "finish_reasons": {
136
+ "stop": 1
137
+ },
138
+ "hit_token_budget": false,
139
+ "hit_token_budget_count": 0,
140
+ "multiplier_vs_ar": 2.560816141204823,
141
+ "quality_passed": true,
142
+ "tok_s": 33.430933263285176,
143
+ "verify_time_s": 36.15608408872504
144
+ },
145
+ {
146
+ "acceptance_by_position": [
147
+ 0.9433962264150944,
148
+ 0.8814016172506739,
149
+ 0.8059299191374663
150
+ ],
151
+ "depth": 3,
152
+ "finish_reasons": {
153
+ "stop": 1
154
+ },
155
+ "hit_token_budget": false,
156
+ "hit_token_budget_count": 0,
157
+ "multiplier_vs_ar": 3.002168329361034,
158
+ "quality_passed": true,
159
+ "tok_s": 39.192696206919734,
160
+ "verify_time_s": 31.090109037584625
161
+ }
162
+ ],
163
+ "greedy_diagnostic": {
164
+ "tok_s": 13.054796369549772
165
+ },
166
+ "quality_rejected": [],
167
+ "tok_s": [
168
+ 39.192696206919734
169
+ ],
170
+ "verdict": "mtp_depth_wins"
171
+ },
172
+ "verified_on": {
173
+ "hardware": "macOS-26.3.1-arm64-arm-64bit-Mach-O",
174
+ "machine_arch": "arm64",
175
+ "macos": "26.3.1",
176
+ "model": "Qwen3.8-27B-MTPLX-Optimized-Quality",
177
+ "timestamp": "2026-08-19T23:08:02-07:00"
178
+ }
179
+ }
preprocessor_config.json ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "size": {
3
+ "longest_edge": 16777216,
4
+ "shortest_edge": 65536
5
+ },
6
+ "patch_size": 16,
7
+ "temporal_patch_size": 2,
8
+ "merge_size": 2,
9
+ "image_mean": [
10
+ 0.5,
11
+ 0.5,
12
+ 0.5
13
+ ],
14
+ "image_std": [
15
+ 0.5,
16
+ 0.5,
17
+ 0.5
18
+ ],
19
+ "processor_class": "Qwen3VLProcessor",
20
+ "image_processor_type": "Qwen2VLImageProcessorFast"
21
+ }
processor_config.json ADDED
@@ -0,0 +1,17 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "image_processor": {
3
+ "image_processor_type": "Qwen2VLImageProcessor",
4
+ "size": {
5
+ "longest_edge": 16777216,
6
+ "shortest_edge": 65536
7
+ }
8
+ },
9
+ "processor_class": "Qwen3VLProcessor",
10
+ "video_processor": {
11
+ "size": {
12
+ "longest_edge": 25165824,
13
+ "shortest_edge": 4096
14
+ },
15
+ "video_processor_type": "Qwen3VLVideoProcessor"
16
+ }
17
+ }
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523
3
+ size 19989325
tokenizer_config.json ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "audio_bos_token": "<|audio_start|>",
4
+ "audio_eos_token": "<|audio_end|>",
5
+ "audio_token": "<|audio_pad|>",
6
+ "backend": "tokenizers",
7
+ "bos_token": null,
8
+ "clean_up_tokenization_spaces": false,
9
+ "eos_token": "<|im_end|>",
10
+ "errors": "replace",
11
+ "image_token": "<|image_pad|>",
12
+ "is_local": true,
13
+ "local_files_only": false,
14
+ "model_max_length": 262144,
15
+ "model_specific_special_tokens": {
16
+ "audio_bos_token": "<|audio_start|>",
17
+ "audio_eos_token": "<|audio_end|>",
18
+ "audio_token": "<|audio_pad|>",
19
+ "image_token": "<|image_pad|>",
20
+ "video_token": "<|video_pad|>",
21
+ "vision_bos_token": "<|vision_start|>",
22
+ "vision_eos_token": "<|vision_end|>"
23
+ },
24
+ "pad_token": "<|endoftext|>",
25
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
26
+ "split_special_tokens": false,
27
+ "tokenizer_class": "Qwen2Tokenizer",
28
+ "tool_parser_type": "qwen3_coder",
29
+ "unk_token": null,
30
+ "video_token": "<|video_pad|>",
31
+ "vision_bos_token": "<|vision_start|>",
32
+ "vision_eos_token": "<|vision_end|>"
33
+ }
video_preprocessor_config.json ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "size": {
3
+ "longest_edge": 25165824,
4
+ "shortest_edge": 4096
5
+ },
6
+ "patch_size": 16,
7
+ "temporal_patch_size": 2,
8
+ "merge_size": 2,
9
+ "image_mean": [
10
+ 0.5,
11
+ 0.5,
12
+ 0.5
13
+ ],
14
+ "image_std": [
15
+ 0.5,
16
+ 0.5,
17
+ 0.5
18
+ ],
19
+ "processor_class": "Qwen3VLProcessor",
20
+ "video_processor_type": "Qwen3VLVideoProcessor"
21
+ }