--- license: other license_name: qwen-community-1.0 license_link: LICENSE library_name: mtplx pipeline_tag: text-generation base_model: Qwen/Qwen3.8-Flash-Next base_model_relation: quantized tags: - mtplx - mlx - apple-silicon - macos - speculative-decoding - multi-token-prediction - qwen - qwen3.8 - qwen3.8-flash-next - flash-next - moe - mtp - 8-bit - vision - mac-studio --- # Qwen 3.8 Flash-Next Optimized Quality **The 8-bit build of Qwen 3.8 Flash-Next, for Macs with 256 GB or 512 GB. Requires MTPLX 2.12.0 or later.** Qwen's 125B-A6B Flash-Next, the Qwen4-generation hybrid mixture of experts with Qwen Sparse Attention and a 51B-parameter n-gram table, packed for [MTPLX](https://mtplx.com) with its multi-token prediction head. The main model and the draft head are 8-bit with group size 64, the structural weights stay in BF16, and the n-gram table is 4-bit with group size 32. On a Mac with 128 GB or more, [Optimized Speed](https://huggingface.co/Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed) is the recommended build. ## Memory The model weights, the draft head and the vision tower need about 128.5 GiB, and the 32 GB n-gram table streams from SSD. - **128 GB**: Cannot load. The weights, the draft head and the vision tower alone need about 128.5 GiB. - **256 GB and 512 GB**: the Macs this pack is for, with about 59.5 GiB left for context and the session cache. The MTPLX app and CLI list it second there, after Optimized Speed. ## Speed Speed on 256 GB and 512 GB Macs is not measured yet. The 8-bit weights move twice the bytes per token of Optimized Speed, so expect slower decoding. ## What is in the pack | Tensor class | Stored precision | Size (GB) | |---|---|---:| | attention | Q8/g64 affine; BF16 scales and biases | 0.635044 | | embeddings | Q8/g64 affine; BF16 scales and biases | 0.675430 | | gdn A log | BF16 | 0.000003 | | gdn conv | BF16 | 0.002949 | | gdn dt bias | BF16 | 0.000003 | | gdn projections | Q8/g64 affine; BF16 scales and biases | 2.215342 | | hyper connections | BF16 | 1.279263 | | lm head | Q8/g64 affine; BF16 scales and biases | 0.675430 | | mtp attention | Q8/g64 affine; BF16 scales and biases | 0.052920 | | mtp fc | BF16 | 0.026214 | | mtp hyper connections | BF16 | 0.039485 | | mtp norms | BF16 | 0.000089 | | mtp qsa indexer | Q8/g64 affine; BF16 scales and biases | 0.001741 | | mtp routed experts | Q8/g64 affine; BF16 scales and biases | 2.673869 | | mtp router | Q8/g64 affine; BF16 scales and biases | 0.001395 | | mtp shared expert | Q8/g64 affine; BF16 scales and biases | 0.005222 | | ngram table | Q4/g32 affine; BF16 scales and biases | 32.000154 | | norms | BF16 | 0.002076 | | ple integer buffers | I64 | 0.000000 | | ple projections and conv | BF16 | 0.065618 | | qsa indexer | Q8/g64 affine; BF16 scales and biases | 0.020890 | | routed experts | Q8/g64 affine; BF16 scales and biases | 128.345702 | | router | Q8/g64 affine; BF16 scales and biases | 0.066977 | | shared expert | Q8/g64 affine; BF16 scales and biases | 0.250675 | | vision tower | BF16 | 0.897862 | The download is 169.96 GB. `size-checksums.json` lists the size and SHA-256 of every other file. ## Use it In the Mac app, pick Qwen 3.8 Flash-Next Optimized Quality. From the command line: ```bash pip install mtplx mtplx serve --model Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality --model-id mtplx-flash-next-optimized-quality ``` MTPLX samples at the official Qwen 3.8 settings (temperature 1.0, top-p 0.95, top-k 20), and drafts are accepted with exact speculative sampling, so the output follows the model's own distribution. Built with the `flash-next-optimized-quality` recipe from [Qwen/Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) at revision `de4b8e4d43b917e7706784d8bb445c9af86a3540`. Qwen Community License, preserved in `LICENSE`. The upstream model card is preserved as `README-upstream-qwen.md`. License and credits are carried from [Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed](https://huggingface.co/Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed). Conversion and serving: MTPLX.