Qwen3.5-9B-MTP-8bit / README.md
ulises-c's picture
Document compatible MLX target models (8bit/6bit)
801e45b verified
|
Raw History Blame Contribute Delete
2.7 kB
---
language:
- en
license: apache-2.0
base_model:
- Qwen/Qwen3.5-9B
- mlx-community/Qwen3.5-9B-MTP-bf16
library_name: mlx
pipeline_tag: text-generation
tags:
- mlx
- mlx-vlm
- qwen3_5_mtp
- qwen3.5
- qwen3.5-9b
- qwen
- mtp
- speculative-decoding
- draft-model
- 8bit
- quantized
inference: false
---
# Qwen3.5-9B-MTP-8bit
This repository contains 8-bit quantized Multi-Token Prediction (MTP) drafter weights for `Qwen/Qwen3.5-9B`, for use with `mlx-vlm` speculative decoding. It was quantized from [`mlx-community/Qwen3.5-9B-MTP-bf16`](https://huggingface.co/mlx-community/Qwen3.5-9B-MTP-bf16).
This is not a standalone chat or text-generation model. Load it as the draft model alongside a compatible Qwen3.5 9B target checkpoint.
## Compatible target models
Pair this drafter with a Qwen3.5-9B target checkpoint. Recommended MLX targets (match target precision to your memory budget):
- [`mlx-community/Qwen3.5-9B-8bit`](https://huggingface.co/mlx-community/Qwen3.5-9B-8bit) — recommended pairing for this 8-bit drafter
- [`mlx-community/Qwen3.5-9B-6bit`](https://huggingface.co/mlx-community/Qwen3.5-9B-6bit)
## Use with mlx-vlm
```bash
uv run mlx_vlm.generate \
--model mlx-community/Qwen3.5-9B-8bit \
--draft-model ulises-c/Qwen3.5-9B-MTP-8bit \
--prompt "Hi, how are you?" \
--max-tokens 256 \
--enable-thinking
```
For local weights:
```bash
uv run mlx_vlm.generate \
--model /path/to/target-model \
--draft-model /path/to/Qwen3.5-9B-MTP-8bit \
--prompt "Hi, how are you?" \
--max-tokens 256 \
--enable-thinking
```
## Model Details
- Model type: `qwen3_5_mtp`
- MTP block size: `2`
- Target architecture: Qwen3.5 9B
- Precision: 8-bit (affine, group size 64) — 8.501 bits per weight
- Runtime: MLX / `mlx-vlm`
- Format: Safetensors with MLX-compatible config and tokenizer files
## Conversion
Quantized with `mlx-vlm` (the `qwen3_5_mtp` architecture is provided by `mlx-vlm`, not `mlx-lm`):
```bash
uvx --from 'mlx-vlm @ git+https://github.com/Blaizzy/mlx-vlm' mlx_vlm.convert \
--hf-path mlx-community/Qwen3.5-9B-MTP-bf16 \
--mlx-path Qwen3.5-9B-MTP-8bit \
-q --q-bits 8 --q-group-size 64
```
## Intended Use
Use this repo only as a speculative decoding drafter for compatible Qwen3.5 9B checkpoints. The target model verifies drafted tokens, while this MTP model proposes candidate tokens per decoding step.
## Limitations
This checkpoint requires runtime support for Qwen/DeepSeek MTP draft models in `mlx-vlm`. Standard standalone generation through generic Transformers APIs is not expected to work with this repository by itself.
Please refer to the upstream `Qwen/Qwen3.5-9B` model card and license terms for model usage constraints.