Image-Text-to-Text
MLX
Safetensors
qwen4_exp
mlx-vlm
omlx
qwen
qwen3.8
mixture-of-experts
vision-language
quantized
apple-silicon
4-bit precision
conversational
Instructions to use TensorFold/Qwen3.8-Flash-Next-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use TensorFold/Qwen3.8-Flash-Next-MLX-4bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("TensorFold/Qwen3.8-Flash-Next-MLX-4bit") config = load_config("TensorFold/Qwen3.8-Flash-Next-MLX-4bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use TensorFold/Qwen3.8-Flash-Next-MLX-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TensorFold/Qwen3.8-Flash-Next-MLX-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "TensorFold/Qwen3.8-Flash-Next-MLX-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use TensorFold/Qwen3.8-Flash-Next-MLX-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TensorFold/Qwen3.8-Flash-Next-MLX-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default TensorFold/Qwen3.8-Flash-Next-MLX-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use TensorFold/Qwen3.8-Flash-Next-MLX-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "TensorFold/Qwen3.8-Flash-Next-MLX-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "TensorFold/Qwen3.8-Flash-Next-MLX-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Add M3 Studio performance and MTP compatibility note
Browse files
README.md
CHANGED
|
@@ -58,6 +58,9 @@ The upstream tokenizer, chat template, vision processor, and generation configur
|
|
| 58 |
> [!IMPORTANT]
|
| 59 |
> Qwen3.8 Flash Next uses the new `qwen4_exp` architecture. Use an oMLX or MLX-VLM build that explicitly lists `qwen4_exp` support. Older MLX-VLM releases cannot load this checkpoint.
|
| 60 |
|
|
|
|
|
|
|
|
|
|
| 61 |
## Quick start
|
| 62 |
|
| 63 |
```bash
|
|
@@ -74,6 +77,18 @@ python -m mlx_vlm.generate \
|
|
| 74 |
--max-tokens 512
|
| 75 |
```
|
| 76 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 77 |
## Architecture
|
| 78 |
|
| 79 |
Qwen3.8 Flash Next is an experimental vision-language architecture combining Gated DeltaNet, Qwen Sparse Attention, sparse mixture-of-experts layers, widened gated residual streams, and hashed bigram/trigram embeddings.
|
|
@@ -95,7 +110,7 @@ For upstream evaluations, intended use, limitations, safety guidance, and the co
|
|
| 95 |
- Source: official BF16 checkpoint.
|
| 96 |
- All 3,671 converted tensors and 22 indexed shards were checked locally.
|
| 97 |
- The release payload was scanned for credentials, personal contact details, private paths, private network information, logs, caches, and private organisation data.
|
| 98 |
-
-
|
| 99 |
- Quantisation can reduce output quality relative to BF16. Test the model on representative workloads before production use.
|
| 100 |
|
| 101 |
This is a community conversion, not an official Qwen release.
|
|
|
|
| 58 |
> [!IMPORTANT]
|
| 59 |
> Qwen3.8 Flash Next uses the new `qwen4_exp` architecture. Use an oMLX or MLX-VLM build that explicitly lists `qwen4_exp` support. Older MLX-VLM releases cannot load this checkpoint.
|
| 60 |
|
| 61 |
+
> [!CAUTION]
|
| 62 |
+
> Do not attach a Qwen3.8 27B MTP drafter to this model. The hidden sizes differ and the drafter is incompatible with Flash Next.
|
| 63 |
+
|
| 64 |
## Quick start
|
| 65 |
|
| 66 |
```bash
|
|
|
|
| 77 |
--max-tokens 512
|
| 78 |
```
|
| 79 |
|
| 80 |
+
## Measured performance
|
| 81 |
+
|
| 82 |
+
Validated on an Apple M3 Studio with text-only generation after model load:
|
| 83 |
+
|
| 84 |
+
| Test path | Result |
|
| 85 |
+
| --- | ---: |
|
| 86 |
+
| oMLX server, warmed 543–566-token responses | 24.1–24.2 tokens/s |
|
| 87 |
+
| oMLX server, warmed shorter responses | 24.6–26.1 tokens/s |
|
| 88 |
+
| Standalone MLX exact-copy smoke test | 31.0 tokens/s |
|
| 89 |
+
|
| 90 |
+
The standalone result is a short smoke test; the longer oMLX figures better represent sustained chat generation. Results vary with prompt length, cache state, sampling settings, runtime version, and memory pressure.
|
| 91 |
+
|
| 92 |
## Architecture
|
| 93 |
|
| 94 |
Qwen3.8 Flash Next is an experimental vision-language architecture combining Gated DeltaNet, Qwen Sparse Attention, sparse mixture-of-experts layers, widened gated residual streams, and hashed bigram/trigram embeddings.
|
|
|
|
| 110 |
- Source: official BF16 checkpoint.
|
| 111 |
- All 3,671 converted tensors and 22 indexed shards were checked locally.
|
| 112 |
- The release payload was scanned for credentials, personal contact details, private paths, private network information, logs, caches, and private organisation data.
|
| 113 |
+
- Deterministic standalone and warmed oMLX server generation tests passed on Apple silicon.
|
| 114 |
- Quantisation can reduce output quality relative to BF16. Test the model on representative workloads before production use.
|
| 115 |
|
| 116 |
This is a community conversion, not an official Qwen release.
|