File size: 5,722 Bytes
a93f2a0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f8b0845
 
 
 
a93f2a0
f8b0845
 
 
 
a93f2a0
 
 
 
 
f8b0845
a93f2a0
f8b0845
 
a93f2a0
 
 
 
 
 
f8b0845
 
 
 
a93f2a0
 
 
 
 
0f4714b
 
ef1c4ff
f8b0845
 
a93f2a0
 
 
 
 
 
 
 
 
 
f8b0845
 
a93f2a0
f8b0845
 
 
a93f2a0
f8b0845
a93f2a0
f8b0845
 
 
a93f2a0
 
 
 
 
 
 
 
 
 
f8b0845
0f4714b
 
f8b0845
a93f2a0
f8b0845
a93f2a0
 
 
 
f8b0845
a93f2a0
 
 
 
 
 
 
 
 
 
f8b0845
 
 
 
 
 
 
a93f2a0
f8b0845
 
 
ef1c4ff
 
 
a93f2a0
f8b0845
a93f2a0
f8b0845
 
 
a93f2a0
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
---
license: apache-2.0
library_name: splash
pipeline_tag: text-generation
inference: false
base_model:
  - mlx-community/Qwen3.6-35B-A3B-4bit
  - incoai/Qwen3.6-35B-A3B-DFlash2
base_model_relation: quantized
tags:
  - splash
  - apple-silicon
  - metal
  - local-inference
  - dflash2
  - speculative-decoding
  - qwen3.6
  - moe
  - 4-bit
---

# Qwen3.6-35B-A3B-Splash

**Qwen3.6-35B-A3B, packed for Splash on Apple silicon.**

[Splash](https://github.com/incoai/splash) is Inco AI's open-source inference
engine for Apple silicon, built around the model. This package contains
everything Splash needs to serve Qwen3.6-35B-A3B: the 4-bit target, its DFlash
2 draft, the vision encoder, and the tokenizer. It is not a Transformers or
MLX checkpoint and does not load anywhere but Splash.

Qwen3.6-35B-A3B is the mixture-of-experts launch model, with 35B total
parameters and about 3B active per token, and the faster of the two. The
other, [Qwen3.8-27B-Splash](https://huggingface.co/incoai/Qwen3.8-27B-Splash),
is a dense 27B model.

[Engine](https://github.com/incoai/splash) 路
[Launch post and benchmarks](https://inco.ai/blog/splash/) 路
[DFlash 2](https://inco.ai/blog/dflash2/)

## Quick start

Apple M3 or newer, macOS 26.4 or later, [Homebrew](https://brew.sh), and 36 GB
of unified memory (48 GB or more recommended).

```bash
brew install incoai/tap/splash
splash serve --model incoai/Qwen3.6-35B-A3B-Splash
```

The first run downloads this package (20.9 GB), verifies it, checks available
memory, and starts serving on `127.0.0.1:8000`. When it prints its `Ready`
line, open <http://127.0.0.1:8000> or attach an agent you already have
installed from another terminal:

```bash
splash opencode    # or: splash claude / splash codex / splash hermes
```

The server binds `127.0.0.1`, and authentication is off by default. Set
`SPLASH_API_KEY` before exposing it beyond your Mac.

The API is OpenAI Chat Completions and Responses, and Anthropic Messages, with
streaming, tool calls, JSON Schema output, images, and inline PDFs:

```bash
curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "incoai/Qwen3.6-35B-A3B-Splash",
    "messages": [{"role": "user", "content": "Explain speculative decoding in one sentence."}]
  }'
```

Reasoning is on by default and is a switch, not a dial. `"reasoning_effort":
"none"` turns it off, and reasoning comes back as `reasoning_content`.

There is no config file. The only settings are ceilings such as `--max-memory`
and `--max-context`, which the
[README](https://github.com/incoai/splash#settings) lists with their defaults.

## Performance

Measured on an M5 Pro (16-core GPU, 48 GB) with selected SPEED-Bench coding
prompts over HTTP, with a 1,024-token output limit and reasoning on. The ratio
in each cell is against the next-fastest engine we measured.

| Metric | Qwen3.6-35B-A3B |
| --- | ---: |
| Decode 路 short prompt | 210 tok/s (1.7脳) |
| Prefill 路 32K prompt | 2,011 tok/s (1.3脳) |
| Time to first token 路 32K prompt, uncached | 17 s (1.3脳) |
| Cached time to first token 路 32K replay | 123 ms (6.6脳) |
| Aggregate decode 路 4 concurrent short prompts | 357 tok/s (2.0脳) |
| Aggregate decode 路 4 concurrent 32K prompts | 236 tok/s (3.8脳) |

Splash led on every measure at every prompt length we tested, and the lead
grows with load. The cached figure replays the prompt exactly, so a real turn
also pays for the tokens it adds. The [launch
post](https://inco.ai/blog/splash/) has the method and the full comparison.

## Package contents

```
target/      42 files   18.2 GiB   Qwen3.6-35B-A3B, 4-bit, one packed file per layer
draft/        7 files    0.5 GiB   DFlash 2 draft model
vision/       1 file     0.8 GiB   bf16 vision encoder
tokenizer/    5 files              tokenizer and chat template
manifest.json                      provenance, geometry, and SHA-256 of every artifact
layout.json                        section-level map of every packed file
```

| Component | Source | Revision |
| --- | --- | --- |
| Target, tokenizer, vision | [`mlx-community/Qwen3.6-35B-A3B-4bit`](https://huggingface.co/mlx-community/Qwen3.6-35B-A3B-4bit) | `38740b847e4cb78f352aba30aa41c76e08e6eb46` |
| Draft | [`incoai/Qwen3.6-35B-A3B-DFlash2`](https://huggingface.co/incoai/Qwen3.6-35B-A3B-DFlash2) | `8e713508f0bb02f03b5cb5cabbc8d9604f924be2` |

The weights are fixed-layout binaries that Splash maps directly from disk,
keeping the upstream conversion's mixed precision: 4-bit weights with 8-bit
expert routers. The draft is a six-layer DFlash 2 model that reads the
target's hidden states at eight layers and proposes 7 tokens per step, which
the target verifies in one pass. The chat template is upstream's with one
change: a system message after the first turn is rendered in place instead of
rejected, which coding agents that inject instructions mid-conversation need.

`manifest.json` records the revisions above, the execution geometry, and the
size and SHA-256 of every artifact. Splash pins an immutable commit of this
repository, checks every artifact's SHA-256 before installing it, and
re-checks sizes and alignment on every start. The target is the upstream 4-bit
conversion, and the target verifies every drafted token, so speculation
changes speed and not the output distribution.

## License

Apache-2.0. Every component is Apache-2.0 upstream as well: Qwen3.6-35B-A3B
(Alibaba), its 4-bit conversion (mlx-community), and the DFlash 2 draft (Inco
AI).

## Citation

```bibtex
@misc{inco2026splash,
  title  = {{Splash: A Local Engine Built Around the Model}},
  author = {{Inco AI}},
  year   = {2026},
  month  = {September},
  url    = {https://inco.ai/blog/splash/}
}
```