File size: 5,722 Bytes
a93f2a0 f8b0845 a93f2a0 f8b0845 a93f2a0 f8b0845 a93f2a0 f8b0845 a93f2a0 f8b0845 a93f2a0 0f4714b ef1c4ff f8b0845 a93f2a0 f8b0845 a93f2a0 f8b0845 a93f2a0 f8b0845 a93f2a0 f8b0845 a93f2a0 f8b0845 0f4714b f8b0845 a93f2a0 f8b0845 a93f2a0 f8b0845 a93f2a0 f8b0845 a93f2a0 f8b0845 ef1c4ff a93f2a0 f8b0845 a93f2a0 f8b0845 a93f2a0 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 | ---
license: apache-2.0
library_name: splash
pipeline_tag: text-generation
inference: false
base_model:
- mlx-community/Qwen3.6-35B-A3B-4bit
- incoai/Qwen3.6-35B-A3B-DFlash2
base_model_relation: quantized
tags:
- splash
- apple-silicon
- metal
- local-inference
- dflash2
- speculative-decoding
- qwen3.6
- moe
- 4-bit
---
# Qwen3.6-35B-A3B-Splash
**Qwen3.6-35B-A3B, packed for Splash on Apple silicon.**
[Splash](https://github.com/incoai/splash) is Inco AI's open-source inference
engine for Apple silicon, built around the model. This package contains
everything Splash needs to serve Qwen3.6-35B-A3B: the 4-bit target, its DFlash
2 draft, the vision encoder, and the tokenizer. It is not a Transformers or
MLX checkpoint and does not load anywhere but Splash.
Qwen3.6-35B-A3B is the mixture-of-experts launch model, with 35B total
parameters and about 3B active per token, and the faster of the two. The
other, [Qwen3.8-27B-Splash](https://huggingface.co/incoai/Qwen3.8-27B-Splash),
is a dense 27B model.
[Engine](https://github.com/incoai/splash) 路
[Launch post and benchmarks](https://inco.ai/blog/splash/) 路
[DFlash 2](https://inco.ai/blog/dflash2/)
## Quick start
Apple M3 or newer, macOS 26.4 or later, [Homebrew](https://brew.sh), and 36 GB
of unified memory (48 GB or more recommended).
```bash
brew install incoai/tap/splash
splash serve --model incoai/Qwen3.6-35B-A3B-Splash
```
The first run downloads this package (20.9 GB), verifies it, checks available
memory, and starts serving on `127.0.0.1:8000`. When it prints its `Ready`
line, open <http://127.0.0.1:8000> or attach an agent you already have
installed from another terminal:
```bash
splash opencode # or: splash claude / splash codex / splash hermes
```
The server binds `127.0.0.1`, and authentication is off by default. Set
`SPLASH_API_KEY` before exposing it beyond your Mac.
The API is OpenAI Chat Completions and Responses, and Anthropic Messages, with
streaming, tool calls, JSON Schema output, images, and inline PDFs:
```bash
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "incoai/Qwen3.6-35B-A3B-Splash",
"messages": [{"role": "user", "content": "Explain speculative decoding in one sentence."}]
}'
```
Reasoning is on by default and is a switch, not a dial. `"reasoning_effort":
"none"` turns it off, and reasoning comes back as `reasoning_content`.
There is no config file. The only settings are ceilings such as `--max-memory`
and `--max-context`, which the
[README](https://github.com/incoai/splash#settings) lists with their defaults.
## Performance
Measured on an M5 Pro (16-core GPU, 48 GB) with selected SPEED-Bench coding
prompts over HTTP, with a 1,024-token output limit and reasoning on. The ratio
in each cell is against the next-fastest engine we measured.
| Metric | Qwen3.6-35B-A3B |
| --- | ---: |
| Decode 路 short prompt | 210 tok/s (1.7脳) |
| Prefill 路 32K prompt | 2,011 tok/s (1.3脳) |
| Time to first token 路 32K prompt, uncached | 17 s (1.3脳) |
| Cached time to first token 路 32K replay | 123 ms (6.6脳) |
| Aggregate decode 路 4 concurrent short prompts | 357 tok/s (2.0脳) |
| Aggregate decode 路 4 concurrent 32K prompts | 236 tok/s (3.8脳) |
Splash led on every measure at every prompt length we tested, and the lead
grows with load. The cached figure replays the prompt exactly, so a real turn
also pays for the tokens it adds. The [launch
post](https://inco.ai/blog/splash/) has the method and the full comparison.
## Package contents
```
target/ 42 files 18.2 GiB Qwen3.6-35B-A3B, 4-bit, one packed file per layer
draft/ 7 files 0.5 GiB DFlash 2 draft model
vision/ 1 file 0.8 GiB bf16 vision encoder
tokenizer/ 5 files tokenizer and chat template
manifest.json provenance, geometry, and SHA-256 of every artifact
layout.json section-level map of every packed file
```
| Component | Source | Revision |
| --- | --- | --- |
| Target, tokenizer, vision | [`mlx-community/Qwen3.6-35B-A3B-4bit`](https://huggingface.co/mlx-community/Qwen3.6-35B-A3B-4bit) | `38740b847e4cb78f352aba30aa41c76e08e6eb46` |
| Draft | [`incoai/Qwen3.6-35B-A3B-DFlash2`](https://huggingface.co/incoai/Qwen3.6-35B-A3B-DFlash2) | `8e713508f0bb02f03b5cb5cabbc8d9604f924be2` |
The weights are fixed-layout binaries that Splash maps directly from disk,
keeping the upstream conversion's mixed precision: 4-bit weights with 8-bit
expert routers. The draft is a six-layer DFlash 2 model that reads the
target's hidden states at eight layers and proposes 7 tokens per step, which
the target verifies in one pass. The chat template is upstream's with one
change: a system message after the first turn is rendered in place instead of
rejected, which coding agents that inject instructions mid-conversation need.
`manifest.json` records the revisions above, the execution geometry, and the
size and SHA-256 of every artifact. Splash pins an immutable commit of this
repository, checks every artifact's SHA-256 before installing it, and
re-checks sizes and alignment on every start. The target is the upstream 4-bit
conversion, and the target verifies every drafted token, so speculation
changes speed and not the output distribution.
## License
Apache-2.0. Every component is Apache-2.0 upstream as well: Qwen3.6-35B-A3B
(Alibaba), its 4-bit conversion (mlx-community), and the DFlash 2 draft (Inco
AI).
## Citation
```bibtex
@misc{inco2026splash,
title = {{Splash: A Local Engine Built Around the Model}},
author = {{Inco AI}},
year = {2026},
month = {September},
url = {https://inco.ai/blog/splash/}
}
```
|