Youssofal commited on
Commit
b8fb374
·
verified ·
1 Parent(s): 3908bbf

Plain card for the 2.12.0 release under the mtplx library: memory, speed status and the tensor table

Browse files
Files changed (2) hide show
  1. README.md +57 -31
  2. size-checksums.json +3 -3
README.md CHANGED
@@ -2,26 +2,56 @@
2
  license: other
3
  license_name: qwen-community-1.0
4
  license_link: LICENSE
5
- library_name: mlx
6
- pipeline_tag: image-text-to-text
7
  base_model: Qwen/Qwen3.8-Flash-Next
8
  base_model_relation: quantized
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9
  ---
10
 
11
- # Qwen3.8-Flash-Next-MTPLX-Optimized-Quality
12
 
13
- > **Requires the upcoming MTPLX 2.12.0 release.**
14
- > The model is published ahead of engine support, which arrives in that release.
15
 
16
- Requires MTPLX 2.12.0 or later. Served id: `mtplx-flash-next-optimized-quality`.
17
- Recipe: `flash-next-optimized-quality`. Source revision: `de4b8e4d43b917e7706784d8bb445c9af86a3540`
18
- (hub-commit).
 
 
 
 
19
 
20
- Verification: **streaming-audited**, mode `streaming`.
21
- The table below is derived from every written safetensors header, including sidecars.
22
- BF16 describes storage; the runtime can use a private Q8 hyper-connection copy during decode.
23
 
24
- | Tensor class | Stored precision | Payload GB |
 
 
 
 
 
 
 
 
 
 
 
 
 
25
  |---|---|---:|
26
  | attention | Q8/g64 affine; BF16 scales and biases | 0.635044 |
27
  | embeddings | Q8/g64 affine; BF16 scales and biases | 0.675430 |
@@ -49,30 +79,26 @@ BF16 describes storage; the runtime can use a private Q8 hyper-connection copy d
49
  | shared expert | Q8/g64 affine; BF16 scales and biases | 0.250675 |
50
  | vision tower | BF16 | 0.897862 |
51
 
52
- Tensor payload: 169.934356 GB (158.264 GiB).
53
- Files before card/manifest generation: 169.958588 GB.
54
- `size-checksums.json` records final file sizes and SHA-256, excluding the manifest itself.
55
-
56
- - **128 GB**: Cannot load: body, MTP and vision alone are approximately 128.46 GiB.
57
- - **192 GB**: Stream the n-gram table; use MTPLX_NGRAM_RESIDENT=0. Validate the context workload on this machine.
58
- - **256 GB**: Fits with the Q4 table resident at 128K (planning estimate; measure the final pack).
59
 
60
- Q8 weights transfer more bytes per parameter than the Speed pack, so bandwidth-limited decode is expected to be slower. Speed is unmeasured until the first run; no speed claim.
61
- Full-model load verified: **false**.
62
- The streaming audit checks stored tensors and sampled dequantization against the source.
63
- It does not prove full-model chat, tool calling, image handling, long-context quality,
64
- or a Speed-pack performance comparison. Those require a run on a supported Mac.
65
 
66
- Recommended memory: **256 GB or 512 GB**. On these Macs, MTPLX's automatic memory
67
- policy normally keeps the n-gram table in RAM as well as the model weights.
68
- The SSD-backed table is used when the memory policy calls for it.
69
 
70
  ```bash
 
71
  mtplx serve --model Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality --model-id mtplx-flash-next-optimized-quality
72
  ```
73
 
74
- Base model and weights: [Qwen/Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next).
75
- Qwen Community License, preserved in `LICENSE`.
76
- Upstream model card preserved as `README-upstream-qwen.md`.
77
- License and credits carried from [Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed](https://huggingface.co/Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed).
 
 
 
 
 
 
78
  Conversion and serving: MTPLX.
 
2
  license: other
3
  license_name: qwen-community-1.0
4
  license_link: LICENSE
5
+ library_name: mtplx
6
+ pipeline_tag: text-generation
7
  base_model: Qwen/Qwen3.8-Flash-Next
8
  base_model_relation: quantized
9
+ tags:
10
+ - mtplx
11
+ - mlx
12
+ - apple-silicon
13
+ - macos
14
+ - speculative-decoding
15
+ - multi-token-prediction
16
+ - qwen
17
+ - qwen3.8
18
+ - qwen3.8-flash-next
19
+ - flash-next
20
+ - moe
21
+ - mtp
22
+ - 8-bit
23
+ - vision
24
+ - mac-studio
25
  ---
26
 
27
+ # Qwen 3.8 Flash-Next Optimized Quality
28
 
29
+ **The 8-bit build of Qwen 3.8 Flash-Next, for Macs with 256 GB or 512 GB. Requires MTPLX 2.12.0 or later.**
 
30
 
31
+ Qwen's 125B-A6B Flash-Next, the Qwen4-generation hybrid mixture of experts with
32
+ Qwen Sparse Attention and a 51B-parameter n-gram table, packed for
33
+ [MTPLX](https://mtplx.com) with its multi-token prediction head. The main model
34
+ and the draft head are 8-bit with group size 64, the structural weights stay in
35
+ BF16, and the n-gram table is 4-bit with group size 32. On a Mac with 128 GB or
36
+ more, [Optimized Speed](https://huggingface.co/Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed) is the
37
+ recommended build.
38
 
39
+ ## Memory
 
 
40
 
41
+ The model weights, the draft head and the vision tower need about 128.5 GiB,
42
+ and the 32 GB n-gram table streams from SSD.
43
+
44
+ - **128 GB**: Cannot load. The weights, the draft head and the vision tower alone need about 128.5 GiB.
45
+ - **256 GB and 512 GB**: the Macs this pack is for, with about 59.5 GiB left for context and the session cache. The MTPLX app and CLI list it second there, after Optimized Speed.
46
+
47
+ ## Speed
48
+
49
+ Speed on 256 GB and 512 GB Macs is not measured yet. The 8-bit weights move
50
+ twice the bytes per token of Optimized Speed, so expect slower decoding.
51
+
52
+ ## What is in the pack
53
+
54
+ | Tensor class | Stored precision | Size (GB) |
55
  |---|---|---:|
56
  | attention | Q8/g64 affine; BF16 scales and biases | 0.635044 |
57
  | embeddings | Q8/g64 affine; BF16 scales and biases | 0.675430 |
 
79
  | shared expert | Q8/g64 affine; BF16 scales and biases | 0.250675 |
80
  | vision tower | BF16 | 0.897862 |
81
 
82
+ The download is 169.96 GB. `size-checksums.json` lists the
83
+ size and SHA-256 of every other file.
 
 
 
 
 
84
 
85
+ ## Use it
 
 
 
 
86
 
87
+ In the Mac app, pick Qwen 3.8 Flash-Next Optimized Quality. From the command line:
 
 
88
 
89
  ```bash
90
+ pip install mtplx
91
  mtplx serve --model Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality --model-id mtplx-flash-next-optimized-quality
92
  ```
93
 
94
+ MTPLX samples at the official Qwen 3.8 settings (temperature 1.0, top-p 0.95,
95
+ top-k 20), and drafts are accepted with exact speculative sampling, so the
96
+ output follows the model's own distribution.
97
+
98
+ Built with the `flash-next-optimized-quality` recipe from
99
+ [Qwen/Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) at
100
+ revision `de4b8e4d43b917e7706784d8bb445c9af86a3540`. Qwen Community License, preserved in
101
+ `LICENSE`. The upstream model card is preserved as
102
+ `README-upstream-qwen.md`. License and credits are carried from
103
+ [Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed](https://huggingface.co/Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed).
104
  Conversion and serving: MTPLX.
size-checksums.json CHANGED
@@ -9,8 +9,8 @@
9
  "sha256": "35ca37ccc366f1ba478dab33841a2c0c18ce53fd62f291ca05341f7728b225b2"
10
  },
11
  "README.md": {
12
- "bytes": 3938,
13
- "sha256": "02ac6e8edf6650ac24c5661a20186e95434df2380c09e76c7bf31b141c391633"
14
  },
15
  "chat_template.jinja": {
16
  "bytes": 8952,
@@ -225,6 +225,6 @@
225
  "sha256": "ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003"
226
  }
227
  },
228
- "total_bytes_excluding_manifest": 169958526962,
229
  "weight_hash_source": "Successful streaming audit from the completed build; immutable weight files reused."
230
  }
 
9
  "sha256": "35ca37ccc366f1ba478dab33841a2c0c18ce53fd62f291ca05341f7728b225b2"
10
  },
11
  "README.md": {
12
+ "bytes": 4101,
13
+ "sha256": "608a2e2b81ff282a871d1f3bec8e76a8612b30665f2726a8aeada0feb91605e6"
14
  },
15
  "chat_template.jinja": {
16
  "bytes": 8952,
 
225
  "sha256": "ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003"
226
  }
227
  },
228
+ "total_bytes_excluding_manifest": 169958527125,
229
  "weight_hash_source": "Successful streaming audit from the completed build; immutable weight files reused."
230
  }