--- language: en tags: - quantized - mlx base_model: - deepseek-ai/DeepSeek-V4-Pro base_model_relation: quantized library_name: mlx pipeline_tag: text-generation --- # DeepSeek-V4-Pro MTP See DeepSeek-V4-Pro with MTP in action: [demonstration videos](https://youtube.com/xcreate) This draft model contains the extracted **Multi-Token Prediction (MTP)** layers to be used alongside the [DeepSeek-V4-Pro-MLX](https://huggingface.co/models?search=inferencerlabs/deepseek-v4-pro-mlx) model as a speculative decoder for improved performance. #### Q2.8 Tested on an M3 Ultra 512 GiB RAM and M4 Max using [Inferencer app](https://inferencer.com)'s distributed compute
| Distributed inference | ~31 tokens/s @ 1000 tokens ~148.5 GiB |
| Text inference (with MTP) | Untested |
Enable speculative decoding in Inference Contols: