--- language: - en license: apache-2.0 base_model: - Qwen/Qwen3.5-9B - mlx-community/Qwen3.5-9B-MTP-bf16 library_name: mlx pipeline_tag: text-generation tags: - mlx - mlx-vlm - qwen3_5_mtp - qwen3.5 - qwen3.5-9b - qwen - mtp - speculative-decoding - draft-model - 8bit - quantized inference: false --- # Qwen3.5-9B-MTP-8bit This repository contains 8-bit quantized Multi-Token Prediction (MTP) drafter weights for `Qwen/Qwen3.5-9B`, for use with `mlx-vlm` speculative decoding. It was quantized from [`mlx-community/Qwen3.5-9B-MTP-bf16`](https://huggingface.co/mlx-community/Qwen3.5-9B-MTP-bf16). This is not a standalone chat or text-generation model. Load it as the draft model alongside a compatible Qwen3.5 9B target checkpoint. ## Compatible target models Pair this drafter with a Qwen3.5-9B target checkpoint. Recommended MLX targets (match target precision to your memory budget): - [`mlx-community/Qwen3.5-9B-8bit`](https://huggingface.co/mlx-community/Qwen3.5-9B-8bit) — recommended pairing for this 8-bit drafter - [`mlx-community/Qwen3.5-9B-6bit`](https://huggingface.co/mlx-community/Qwen3.5-9B-6bit) ## Use with mlx-vlm ```bash uv run mlx_vlm.generate \ --model mlx-community/Qwen3.5-9B-8bit \ --draft-model ulises-c/Qwen3.5-9B-MTP-8bit \ --prompt "Hi, how are you?" \ --max-tokens 256 \ --enable-thinking ``` For local weights: ```bash uv run mlx_vlm.generate \ --model /path/to/target-model \ --draft-model /path/to/Qwen3.5-9B-MTP-8bit \ --prompt "Hi, how are you?" \ --max-tokens 256 \ --enable-thinking ``` ## Model Details - Model type: `qwen3_5_mtp` - MTP block size: `2` - Target architecture: Qwen3.5 9B - Precision: 8-bit (affine, group size 64) — 8.501 bits per weight - Runtime: MLX / `mlx-vlm` - Format: Safetensors with MLX-compatible config and tokenizer files ## Conversion Quantized with `mlx-vlm` (the `qwen3_5_mtp` architecture is provided by `mlx-vlm`, not `mlx-lm`): ```bash uvx --from 'mlx-vlm @ git+https://github.com/Blaizzy/mlx-vlm' mlx_vlm.convert \ --hf-path mlx-community/Qwen3.5-9B-MTP-bf16 \ --mlx-path Qwen3.5-9B-MTP-8bit \ -q --q-bits 8 --q-group-size 64 ``` ## Intended Use Use this repo only as a speculative decoding drafter for compatible Qwen3.5 9B checkpoints. The target model verifies drafted tokens, while this MTP model proposes candidate tokens per decoding step. ## Limitations This checkpoint requires runtime support for Qwen/DeepSeek MTP draft models in `mlx-vlm`. Standard standalone generation through generic Transformers APIs is not expected to work with this repository by itself. Please refer to the upstream `Qwen/Qwen3.5-9B` model card and license terms for model usage constraints.