--- language: en tags: - mlx base_model: - deepseek-ai/DeepSeek-V4-Flash-0731 library_name: mlx pipeline_tag: text-generation --- # DeepSeek-V4-Flash-0731 DSpark See DeepSeek-V4-Flash with MTP in action: [demonstration videos](https://youtube.com/xcreate) This draft model contains the extracted **DSpark** layers to be used alongside the [DeepSeek-V4-Flash-0731-MLX](https://huggingface.co/models?search=inferencerlabs/deepseek-v4-flash-0731-mlx) model as a speculative decoder for improved performance. #### Tested on an M3 Ultra 512 GiB RAM using [Inferencer app v2.3.1](https://inferencer.com) ### Tetris HTML
| Text inference | ~31 tokens/s @ 500 tokens ~148.5 GiB |
| Text inference (with MTP) | ~35.8 tokens/s @ 500 tokens ~152.9 GiB |
| Text inference (with DSpark) | ~53.1 tokens/s @ 500 tokens ~165.6 GiB |
| Text inference | ~30.5 tokens/s @ 500 tokens ~148.5 GiB |
| Text inference (with MTP) | ~37.3 tokens/s @ 500 tokens ~152.9 GiB |
| Text inference (with DSpark) | ~31.3 tokens/s @ 500 tokens ~165.6 GiB |
Enable speculative decoding in Inference Contols: