|
Download README.md from incoai/Qwen3.6-35B-A3B-Splash: direct link, hf CLI and curl.
- Browser
- Download file 5.72 kB
-
https://huggingface.co/incoai/Qwen3.6-35B-A3B-Splash/resolve/main/README.md
- Command line
-
hf download hf://incoai/Qwen3.6-35B-A3B-Splash/README.md
-
curl -L -o README.md https://huggingface.co/incoai/Qwen3.6-35B-A3B-Splash/resolve/main/README.md
5.72 kB
| license: apache-2.0 | |
| library_name: splash | |
| pipeline_tag: text-generation | |
| inference: false | |
| base_model: | |
| - mlx-community/Qwen3.6-35B-A3B-4bit | |
| - incoai/Qwen3.6-35B-A3B-DFlash2 | |
| base_model_relation: quantized | |
| tags: | |
| - splash | |
| - apple-silicon | |
| - metal | |
| - local-inference | |
| - dflash2 | |
| - speculative-decoding | |
| - qwen3.6 | |
| - moe | |
| - 4-bit | |
| # Qwen3.6-35B-A3B-Splash | |
| **Qwen3.6-35B-A3B, packed for Splash on Apple silicon.** | |
| [Splash](https://github.com/incoai/splash) is Inco AI's open-source inference | |
| engine for Apple silicon, built around the model. This package contains | |
| everything Splash needs to serve Qwen3.6-35B-A3B: the 4-bit target, its DFlash | |
| 2 draft, the vision encoder, and the tokenizer. It is not a Transformers or | |
| MLX checkpoint and does not load anywhere but Splash. | |
| Qwen3.6-35B-A3B is the mixture-of-experts launch model, with 35B total | |
| parameters and about 3B active per token, and the faster of the two. The | |
| other, [Qwen3.8-27B-Splash](https://huggingface.co/incoai/Qwen3.8-27B-Splash), | |
| is a dense 27B model. | |
| [Engine](https://github.com/incoai/splash) 路 | |
| [Launch post and benchmarks](https://inco.ai/blog/splash/) 路 | |
| [DFlash 2](https://inco.ai/blog/dflash2/) | |
| ## Quick start | |
| Apple M3 or newer, macOS 26.4 or later, [Homebrew](https://brew.sh), and 36 GB | |
| of unified memory (48 GB or more recommended). | |
| ```bash | |
| brew install incoai/tap/splash | |
| splash serve --model incoai/Qwen3.6-35B-A3B-Splash | |
| ``` | |
| The first run downloads this package (20.9 GB), verifies it, checks available | |
| memory, and starts serving on `127.0.0.1:8000`. When it prints its `Ready` | |
| line, open <http://127.0.0.1:8000> or attach an agent you already have | |
| installed from another terminal: | |
| ```bash | |
| splash opencode # or: splash claude / splash codex / splash hermes | |
| ``` | |
| The server binds `127.0.0.1`, and authentication is off by default. Set | |
| `SPLASH_API_KEY` before exposing it beyond your Mac. | |
| The API is OpenAI Chat Completions and Responses, and Anthropic Messages, with | |
| streaming, tool calls, JSON Schema output, images, and inline PDFs: | |
| ```bash | |
| curl http://127.0.0.1:8000/v1/chat/completions \ | |
| -H 'Content-Type: application/json' \ | |
| -d '{ | |
| "model": "incoai/Qwen3.6-35B-A3B-Splash", | |
| "messages": [{"role": "user", "content": "Explain speculative decoding in one sentence."}] | |
| }' | |
| ``` | |
| Reasoning is on by default and is a switch, not a dial. `"reasoning_effort": | |
| "none"` turns it off, and reasoning comes back as `reasoning_content`. | |
| There is no config file. The only settings are ceilings such as `--max-memory` | |
| and `--max-context`, which the | |
| [README](https://github.com/incoai/splash#settings) lists with their defaults. | |
| ## Performance | |
| Measured on an M5 Pro (16-core GPU, 48 GB) with selected SPEED-Bench coding | |
| prompts over HTTP, with a 1,024-token output limit and reasoning on. The ratio | |
| in each cell is against the next-fastest engine we measured. | |
| | Metric | Qwen3.6-35B-A3B | | |
| | --- | ---: | | |
| | Decode 路 short prompt | 210 tok/s (1.7脳) | | |
| | Prefill 路 32K prompt | 2,011 tok/s (1.3脳) | | |
| | Time to first token 路 32K prompt, uncached | 17 s (1.3脳) | | |
| | Cached time to first token 路 32K replay | 123 ms (6.6脳) | | |
| | Aggregate decode 路 4 concurrent short prompts | 357 tok/s (2.0脳) | | |
| | Aggregate decode 路 4 concurrent 32K prompts | 236 tok/s (3.8脳) | | |
| Splash led on every measure at every prompt length we tested, and the lead | |
| grows with load. The cached figure replays the prompt exactly, so a real turn | |
| also pays for the tokens it adds. The [launch | |
| post](https://inco.ai/blog/splash/) has the method and the full comparison. | |
| ## Package contents | |
| ``` | |
| target/ 42 files 18.2 GiB Qwen3.6-35B-A3B, 4-bit, one packed file per layer | |
| draft/ 7 files 0.5 GiB DFlash 2 draft model | |
| vision/ 1 file 0.8 GiB bf16 vision encoder | |
| tokenizer/ 5 files tokenizer and chat template | |
| manifest.json provenance, geometry, and SHA-256 of every artifact | |
| layout.json section-level map of every packed file | |
| ``` | |
| | Component | Source | Revision | | |
| | --- | --- | --- | | |
| | Target, tokenizer, vision | [`mlx-community/Qwen3.6-35B-A3B-4bit`](https://huggingface.co/mlx-community/Qwen3.6-35B-A3B-4bit) | `38740b847e4cb78f352aba30aa41c76e08e6eb46` | | |
| | Draft | [`incoai/Qwen3.6-35B-A3B-DFlash2`](https://huggingface.co/incoai/Qwen3.6-35B-A3B-DFlash2) | `8e713508f0bb02f03b5cb5cabbc8d9604f924be2` | | |
| The weights are fixed-layout binaries that Splash maps directly from disk, | |
| keeping the upstream conversion's mixed precision: 4-bit weights with 8-bit | |
| expert routers. The draft is a six-layer DFlash 2 model that reads the | |
| target's hidden states at eight layers and proposes 7 tokens per step, which | |
| the target verifies in one pass. The chat template is upstream's with one | |
| change: a system message after the first turn is rendered in place instead of | |
| rejected, which coding agents that inject instructions mid-conversation need. | |
| `manifest.json` records the revisions above, the execution geometry, and the | |
| size and SHA-256 of every artifact. Splash pins an immutable commit of this | |
| repository, checks every artifact's SHA-256 before installing it, and | |
| re-checks sizes and alignment on every start. The target is the upstream 4-bit | |
| conversion, and the target verifies every drafted token, so speculation | |
| changes speed and not the output distribution. | |
| ## License | |
| Apache-2.0. Every component is Apache-2.0 upstream as well: Qwen3.6-35B-A3B | |
| (Alibaba), its 4-bit conversion (mlx-community), and the DFlash 2 draft (Inco | |
| AI). | |
| ## Citation | |
| ```bibtex | |
| @misc{inco2026splash, | |
| title = {{Splash: A Local Engine Built Around the Model}}, | |
| author = {{Inco AI}}, | |
| year = {2026}, | |
| month = {September}, | |
| url = {https://inco.ai/blog/splash/} | |
| } | |
| ``` | |