zhijianliu's picture
Model card: authentication is optional via SPLASH_API_KEY
0f4714b verified
|
Raw History Blame Contribute Delete
5.72 kB
---
license: apache-2.0
library_name: splash
pipeline_tag: text-generation
inference: false
base_model:
- mlx-community/Qwen3.6-35B-A3B-4bit
- incoai/Qwen3.6-35B-A3B-DFlash2
base_model_relation: quantized
tags:
- splash
- apple-silicon
- metal
- local-inference
- dflash2
- speculative-decoding
- qwen3.6
- moe
- 4-bit
---
# Qwen3.6-35B-A3B-Splash
**Qwen3.6-35B-A3B, packed for Splash on Apple silicon.**
[Splash](https://github.com/incoai/splash) is Inco AI's open-source inference
engine for Apple silicon, built around the model. This package contains
everything Splash needs to serve Qwen3.6-35B-A3B: the 4-bit target, its DFlash
2 draft, the vision encoder, and the tokenizer. It is not a Transformers or
MLX checkpoint and does not load anywhere but Splash.
Qwen3.6-35B-A3B is the mixture-of-experts launch model, with 35B total
parameters and about 3B active per token, and the faster of the two. The
other, [Qwen3.8-27B-Splash](https://huggingface.co/incoai/Qwen3.8-27B-Splash),
is a dense 27B model.
[Engine](https://github.com/incoai/splash) 路
[Launch post and benchmarks](https://inco.ai/blog/splash/) 路
[DFlash 2](https://inco.ai/blog/dflash2/)
## Quick start
Apple M3 or newer, macOS 26.4 or later, [Homebrew](https://brew.sh), and 36 GB
of unified memory (48 GB or more recommended).
```bash
brew install incoai/tap/splash
splash serve --model incoai/Qwen3.6-35B-A3B-Splash
```
The first run downloads this package (20.9 GB), verifies it, checks available
memory, and starts serving on `127.0.0.1:8000`. When it prints its `Ready`
line, open <http://127.0.0.1:8000> or attach an agent you already have
installed from another terminal:
```bash
splash opencode # or: splash claude / splash codex / splash hermes
```
The server binds `127.0.0.1`, and authentication is off by default. Set
`SPLASH_API_KEY` before exposing it beyond your Mac.
The API is OpenAI Chat Completions and Responses, and Anthropic Messages, with
streaming, tool calls, JSON Schema output, images, and inline PDFs:
```bash
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "incoai/Qwen3.6-35B-A3B-Splash",
"messages": [{"role": "user", "content": "Explain speculative decoding in one sentence."}]
}'
```
Reasoning is on by default and is a switch, not a dial. `"reasoning_effort":
"none"` turns it off, and reasoning comes back as `reasoning_content`.
There is no config file. The only settings are ceilings such as `--max-memory`
and `--max-context`, which the
[README](https://github.com/incoai/splash#settings) lists with their defaults.
## Performance
Measured on an M5 Pro (16-core GPU, 48 GB) with selected SPEED-Bench coding
prompts over HTTP, with a 1,024-token output limit and reasoning on. The ratio
in each cell is against the next-fastest engine we measured.
| Metric | Qwen3.6-35B-A3B |
| --- | ---: |
| Decode 路 short prompt | 210 tok/s (1.7脳) |
| Prefill 路 32K prompt | 2,011 tok/s (1.3脳) |
| Time to first token 路 32K prompt, uncached | 17 s (1.3脳) |
| Cached time to first token 路 32K replay | 123 ms (6.6脳) |
| Aggregate decode 路 4 concurrent short prompts | 357 tok/s (2.0脳) |
| Aggregate decode 路 4 concurrent 32K prompts | 236 tok/s (3.8脳) |
Splash led on every measure at every prompt length we tested, and the lead
grows with load. The cached figure replays the prompt exactly, so a real turn
also pays for the tokens it adds. The [launch
post](https://inco.ai/blog/splash/) has the method and the full comparison.
## Package contents
```
target/ 42 files 18.2 GiB Qwen3.6-35B-A3B, 4-bit, one packed file per layer
draft/ 7 files 0.5 GiB DFlash 2 draft model
vision/ 1 file 0.8 GiB bf16 vision encoder
tokenizer/ 5 files tokenizer and chat template
manifest.json provenance, geometry, and SHA-256 of every artifact
layout.json section-level map of every packed file
```
| Component | Source | Revision |
| --- | --- | --- |
| Target, tokenizer, vision | [`mlx-community/Qwen3.6-35B-A3B-4bit`](https://huggingface.co/mlx-community/Qwen3.6-35B-A3B-4bit) | `38740b847e4cb78f352aba30aa41c76e08e6eb46` |
| Draft | [`incoai/Qwen3.6-35B-A3B-DFlash2`](https://huggingface.co/incoai/Qwen3.6-35B-A3B-DFlash2) | `8e713508f0bb02f03b5cb5cabbc8d9604f924be2` |
The weights are fixed-layout binaries that Splash maps directly from disk,
keeping the upstream conversion's mixed precision: 4-bit weights with 8-bit
expert routers. The draft is a six-layer DFlash 2 model that reads the
target's hidden states at eight layers and proposes 7 tokens per step, which
the target verifies in one pass. The chat template is upstream's with one
change: a system message after the first turn is rendered in place instead of
rejected, which coding agents that inject instructions mid-conversation need.
`manifest.json` records the revisions above, the execution geometry, and the
size and SHA-256 of every artifact. Splash pins an immutable commit of this
repository, checks every artifact's SHA-256 before installing it, and
re-checks sizes and alignment on every start. The target is the upstream 4-bit
conversion, and the target verifies every drafted token, so speculation
changes speed and not the output distribution.
## License
Apache-2.0. Every component is Apache-2.0 upstream as well: Qwen3.6-35B-A3B
(Alibaba), its 4-bit conversion (mlx-community), and the DFlash 2 draft (Inco
AI).
## Citation
```bibtex
@misc{inco2026splash,
title = {{Splash: A Local Engine Built Around the Model}},
author = {{Inco AI}},
year = {2026},
month = {September},
url = {https://inco.ai/blog/splash/}
}
```