Qwen3.6-35B-A3B-Splash (mirror)
⚠️ Settings: read this before you run it
| Setting | Use this | Why |
|---|---|---|
temperature |
0.6 | Qwen's thinking-mode value. Don't go greedy (0): Qwen's guidance is that it makes endless repetition more likely. |
top_p |
0.95 | Honored by the runtime. |
top_k |
20 | Honored. Must be an integer from 1 to 32. Anything higher, or "off", returns HTTP 400: Splash top-k must be an integer from 1 to 32. |
max_tokens |
cap it, ~16k to 20k | This is your only loop guard. Without a cap, one runaway reasoning turn ran 112,000 tokens in about 6.5 minutes and never answered. |
repeat_penalty, presence_penalty, frequency_penalty |
don't rely on them | The runtime ignores all three (byte-identical output with and without them). Setting them does nothing. |
| Thinking | on by default | /no_think doesn't stop it reasoning. The reasoning just goes to reasoning_content. |
LM Studio per-model default (saved as ~/.lmstudio/.internal/user-concrete-model-default-config/incoai/Qwen3.6-35B-A3B-Splash.json, which LM Studio reads live):
{
"preset": "",
"operation": {
"fields": [
{ "key": "llm.prediction.temperature", "value": 0.6 },
{ "key": "llm.prediction.topPSampling", "value": { "checked": true, "value": 0.95 } },
{ "key": "llm.prediction.topKSampling", "value": 20 },
{ "key": "llm.prediction.contextOverflowPolicy", "value": "rollingWindow" }
]
},
"load": { "fields": [] }
}
Every value above is honored except the ones in the "don't rely on them" row. Tested 2026-09-23 against LM Studio's Splash runtime.
A byte-for-byte mirror of incoai/Qwen3.6-35B-A3B-Splash, kept here as a pinned copy for my local agent stack. All credit for the packing, the DFlash 2 draft, and the Splash engine goes to Inco AI. Read their card first. This one only adds what I found running it.
Verified identical to upstream on 2026-09-23: the artifact_set_sha256 in manifest.json matches incoai's published manifest (0c465f96…2440a).
What's in the package
It's not a Transformers or MLX checkpoint. It only loads in Splash, or in LM Studio's Splash runtime.
| Path | Contents |
|---|---|
target/ |
Qwen3.6-35B-A3B, 4-bit packed (splash-packed-q4-moe), 40 layers, 256 experts, 8 active per token |
draft/ |
DFlash 2 draft model, 6 layers, proposes 7 tokens per step |
vision/ |
Vision encoder |
tokenizer/ |
Tokenizer and chat template |
manifest.json, layout.json |
Artifact hashes, execution geometry, upstream revisions |
Upstream sources, from the manifest: target, tokenizer, and vision from mlx-community/Qwen3.6-35B-A3B-4bit @ 38740b8, draft from incoai/Qwen3.6-35B-A3B-DFlash2 @ 8e71350. About 20.9 GB on disk.
Running it
LM Studio: it shows up as qwen3.6-35b-a3b-splash and reports model format yuzu. This is how I run it.
Splash: upstream's command is splash serve --model incoai/Qwen3.6-35B-A3B-Splash. I haven't tried pointing Splash at this mirror's repo id, so use upstream's if you're going that route.
Measured
LM Studio on an M5 Max (128 GB), OpenAI-compatible endpoint, non-streaming: 1,000 tokens in 6.75 s, about 148 tok/s end to end with prefill included. Decode alone runs higher, around 170 tok/s on the same machine. For comparison, the Qwen3.8-27B MTPLX build I was running before did about 24 tok/s in the same agent.
Sampling notes (LM Studio's Splash runtime, 2026-09-23)
I tested each parameter directly against the endpoint:
- Honored:
temperature,top_p,top_k. Also the per-model defaults in LM Studio's model config, which are picked up live. - Ignored:
repeat_penalty,presence_penalty,frequency_penalty. Output at temperature 0 was byte-identical with and without them.
So the usual repetition guard doesn't exist on this runtime (at least not yet). In my first evening of use it went into a reasoning loop: about 112k tokens of the same closing line, repeated inside the think block, and it never emitted an answer. At this speed that's six and a half minutes before anyone notices. If you run it behind an agent, cap max_tokens so a loop dies on its own. I'm running temperature 0.6, top_p 0.95, top_k 20, which is Qwen's thinking-mode recommendation. Don't go greedy, since Qwen's own guidance is that greedy decoding makes endless repetition more likely.
If you drive LM Studio from code
@lmstudio/sdk 1.5.0 hangs (it never resolves or rejects) on llm.listLoaded() whenever a Splash model is loaded, because its schema doesn't know the yuzu format. 2.0.0 fixes it. Upgrade before loading this.
License
Apache 2.0, same as upstream and the Qwen base model.
Model tree for tokenfires/Qwen3.6-35B-A3B-Splash
Base model
Qwen/Qwen3.6-35B-A3B