Image-Text-to-Text
Safetensors
MLX
mtplx
qwen4_exp
apple-silicon
macos
speculative-decoding
multi-token-prediction
qwen
qwen3.8-flash-next
Mixture of Experts
mtp
local-ai
chat
qwen3.8
qwen3-8
qwen-3.8
local-llm
llm
8-bit precision
vision
m5-max
m3-ultra
mac-studio
opencode
claude-code
flash-next
qwen3-8-flash-next
qwen4
125b
conversational
Instructions to use Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality") config = load_config("Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Plain card for the 2.12.0 release under the mtplx library: memory, speed status and the tensor table
Browse files- README.md +57 -31
- size-checksums.json +3 -3
README.md
CHANGED
|
@@ -2,26 +2,56 @@
|
|
| 2 |
license: other
|
| 3 |
license_name: qwen-community-1.0
|
| 4 |
license_link: LICENSE
|
| 5 |
-
library_name:
|
| 6 |
-
pipeline_tag:
|
| 7 |
base_model: Qwen/Qwen3.8-Flash-Next
|
| 8 |
base_model_relation: quantized
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 9 |
---
|
| 10 |
|
| 11 |
-
#
|
| 12 |
|
| 13 |
-
|
| 14 |
-
> The model is published ahead of engine support, which arrives in that release.
|
| 15 |
|
| 16 |
-
|
| 17 |
-
|
| 18 |
-
(
|
|
|
|
|
|
|
|
|
|
|
|
|
| 19 |
|
| 20 |
-
|
| 21 |
-
The table below is derived from every written safetensors header, including sidecars.
|
| 22 |
-
BF16 describes storage; the runtime can use a private Q8 hyper-connection copy during decode.
|
| 23 |
|
| 24 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
|---|---|---:|
|
| 26 |
| attention | Q8/g64 affine; BF16 scales and biases | 0.635044 |
|
| 27 |
| embeddings | Q8/g64 affine; BF16 scales and biases | 0.675430 |
|
|
@@ -49,30 +79,26 @@ BF16 describes storage; the runtime can use a private Q8 hyper-connection copy d
|
|
| 49 |
| shared expert | Q8/g64 affine; BF16 scales and biases | 0.250675 |
|
| 50 |
| vision tower | BF16 | 0.897862 |
|
| 51 |
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
`size-checksums.json` records final file sizes and SHA-256, excluding the manifest itself.
|
| 55 |
-
|
| 56 |
-
- **128 GB**: Cannot load: body, MTP and vision alone are approximately 128.46 GiB.
|
| 57 |
-
- **192 GB**: Stream the n-gram table; use MTPLX_NGRAM_RESIDENT=0. Validate the context workload on this machine.
|
| 58 |
-
- **256 GB**: Fits with the Q4 table resident at 128K (planning estimate; measure the final pack).
|
| 59 |
|
| 60 |
-
|
| 61 |
-
Full-model load verified: **false**.
|
| 62 |
-
The streaming audit checks stored tensors and sampled dequantization against the source.
|
| 63 |
-
It does not prove full-model chat, tool calling, image handling, long-context quality,
|
| 64 |
-
or a Speed-pack performance comparison. Those require a run on a supported Mac.
|
| 65 |
|
| 66 |
-
|
| 67 |
-
policy normally keeps the n-gram table in RAM as well as the model weights.
|
| 68 |
-
The SSD-backed table is used when the memory policy calls for it.
|
| 69 |
|
| 70 |
```bash
|
|
|
|
| 71 |
mtplx serve --model Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality --model-id mtplx-flash-next-optimized-quality
|
| 72 |
```
|
| 73 |
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 78 |
Conversion and serving: MTPLX.
|
|
|
|
| 2 |
license: other
|
| 3 |
license_name: qwen-community-1.0
|
| 4 |
license_link: LICENSE
|
| 5 |
+
library_name: mtplx
|
| 6 |
+
pipeline_tag: text-generation
|
| 7 |
base_model: Qwen/Qwen3.8-Flash-Next
|
| 8 |
base_model_relation: quantized
|
| 9 |
+
tags:
|
| 10 |
+
- mtplx
|
| 11 |
+
- mlx
|
| 12 |
+
- apple-silicon
|
| 13 |
+
- macos
|
| 14 |
+
- speculative-decoding
|
| 15 |
+
- multi-token-prediction
|
| 16 |
+
- qwen
|
| 17 |
+
- qwen3.8
|
| 18 |
+
- qwen3.8-flash-next
|
| 19 |
+
- flash-next
|
| 20 |
+
- moe
|
| 21 |
+
- mtp
|
| 22 |
+
- 8-bit
|
| 23 |
+
- vision
|
| 24 |
+
- mac-studio
|
| 25 |
---
|
| 26 |
|
| 27 |
+
# Qwen 3.8 Flash-Next Optimized Quality
|
| 28 |
|
| 29 |
+
**The 8-bit build of Qwen 3.8 Flash-Next, for Macs with 256 GB or 512 GB. Requires MTPLX 2.12.0 or later.**
|
|
|
|
| 30 |
|
| 31 |
+
Qwen's 125B-A6B Flash-Next, the Qwen4-generation hybrid mixture of experts with
|
| 32 |
+
Qwen Sparse Attention and a 51B-parameter n-gram table, packed for
|
| 33 |
+
[MTPLX](https://mtplx.com) with its multi-token prediction head. The main model
|
| 34 |
+
and the draft head are 8-bit with group size 64, the structural weights stay in
|
| 35 |
+
BF16, and the n-gram table is 4-bit with group size 32. On a Mac with 128 GB or
|
| 36 |
+
more, [Optimized Speed](https://huggingface.co/Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed) is the
|
| 37 |
+
recommended build.
|
| 38 |
|
| 39 |
+
## Memory
|
|
|
|
|
|
|
| 40 |
|
| 41 |
+
The model weights, the draft head and the vision tower need about 128.5 GiB,
|
| 42 |
+
and the 32 GB n-gram table streams from SSD.
|
| 43 |
+
|
| 44 |
+
- **128 GB**: Cannot load. The weights, the draft head and the vision tower alone need about 128.5 GiB.
|
| 45 |
+
- **256 GB and 512 GB**: the Macs this pack is for, with about 59.5 GiB left for context and the session cache. The MTPLX app and CLI list it second there, after Optimized Speed.
|
| 46 |
+
|
| 47 |
+
## Speed
|
| 48 |
+
|
| 49 |
+
Speed on 256 GB and 512 GB Macs is not measured yet. The 8-bit weights move
|
| 50 |
+
twice the bytes per token of Optimized Speed, so expect slower decoding.
|
| 51 |
+
|
| 52 |
+
## What is in the pack
|
| 53 |
+
|
| 54 |
+
| Tensor class | Stored precision | Size (GB) |
|
| 55 |
|---|---|---:|
|
| 56 |
| attention | Q8/g64 affine; BF16 scales and biases | 0.635044 |
|
| 57 |
| embeddings | Q8/g64 affine; BF16 scales and biases | 0.675430 |
|
|
|
|
| 79 |
| shared expert | Q8/g64 affine; BF16 scales and biases | 0.250675 |
|
| 80 |
| vision tower | BF16 | 0.897862 |
|
| 81 |
|
| 82 |
+
The download is 169.96 GB. `size-checksums.json` lists the
|
| 83 |
+
size and SHA-256 of every other file.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 84 |
|
| 85 |
+
## Use it
|
|
|
|
|
|
|
|
|
|
|
|
|
| 86 |
|
| 87 |
+
In the Mac app, pick Qwen 3.8 Flash-Next Optimized Quality. From the command line:
|
|
|
|
|
|
|
| 88 |
|
| 89 |
```bash
|
| 90 |
+
pip install mtplx
|
| 91 |
mtplx serve --model Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Quality --model-id mtplx-flash-next-optimized-quality
|
| 92 |
```
|
| 93 |
|
| 94 |
+
MTPLX samples at the official Qwen 3.8 settings (temperature 1.0, top-p 0.95,
|
| 95 |
+
top-k 20), and drafts are accepted with exact speculative sampling, so the
|
| 96 |
+
output follows the model's own distribution.
|
| 97 |
+
|
| 98 |
+
Built with the `flash-next-optimized-quality` recipe from
|
| 99 |
+
[Qwen/Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) at
|
| 100 |
+
revision `de4b8e4d43b917e7706784d8bb445c9af86a3540`. Qwen Community License, preserved in
|
| 101 |
+
`LICENSE`. The upstream model card is preserved as
|
| 102 |
+
`README-upstream-qwen.md`. License and credits are carried from
|
| 103 |
+
[Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed](https://huggingface.co/Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed).
|
| 104 |
Conversion and serving: MTPLX.
|
size-checksums.json
CHANGED
|
@@ -9,8 +9,8 @@
|
|
| 9 |
"sha256": "35ca37ccc366f1ba478dab33841a2c0c18ce53fd62f291ca05341f7728b225b2"
|
| 10 |
},
|
| 11 |
"README.md": {
|
| 12 |
-
"bytes":
|
| 13 |
-
"sha256": "
|
| 14 |
},
|
| 15 |
"chat_template.jinja": {
|
| 16 |
"bytes": 8952,
|
|
@@ -225,6 +225,6 @@
|
|
| 225 |
"sha256": "ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003"
|
| 226 |
}
|
| 227 |
},
|
| 228 |
-
"total_bytes_excluding_manifest":
|
| 229 |
"weight_hash_source": "Successful streaming audit from the completed build; immutable weight files reused."
|
| 230 |
}
|
|
|
|
| 9 |
"sha256": "35ca37ccc366f1ba478dab33841a2c0c18ce53fd62f291ca05341f7728b225b2"
|
| 10 |
},
|
| 11 |
"README.md": {
|
| 12 |
+
"bytes": 4101,
|
| 13 |
+
"sha256": "608a2e2b81ff282a871d1f3bec8e76a8612b30665f2726a8aeada0feb91605e6"
|
| 14 |
},
|
| 15 |
"chat_template.jinja": {
|
| 16 |
"bytes": 8952,
|
|
|
|
| 225 |
"sha256": "ce99b4cb2983d118806ce0a8b777a35b093e2000a503ebde25853284c9dfa003"
|
| 226 |
}
|
| 227 |
},
|
| 228 |
+
"total_bytes_excluding_manifest": 169958527125,
|
| 229 |
"weight_hash_source": "Successful streaming audit from the completed build; immutable weight files reused."
|
| 230 |
}
|