Qwen3.6-35B-A3B-Heretic-Splash

A community Splash-format conversion of the existing Qwen3.6-35B-A3B Heretic 4-bit checkpoint. This is not newly trained weights, a new Heretic modification, or an official Inco AI/Qwen release.

The existing mixed affine 4-bit weights and 8-bit routers were repacked into Splash's native containers without additional requantization. The stock Inco DFlash2 draft is retained unchanged; it was not retrained for Heretic.

This package is approximately 20.95 GB of runtime artifacts. It is for Splash, not a Transformers, GGUF, or MLX loader.

Run with Splash

Install Splash. Its upstream requirements are Apple M3 or newer, macOS 26.4 or later, and at least 36 GB unified memory; 48 GB or more is recommended upstream. This conversion was tested on an M5 Max with 128 GiB, not on the minimum-memory configuration.

Run:

brew install incoai/tap/splash
splash serve --model tcstacks/Qwen3.6-35B-A3B-Heretic-Splash --max-context 64K

Splash downloads the manifest-listed artifacts, verifies them, and serves on http://127.0.0.1:8000. The standard launcher requires that port to be free; do not stop someone else's service. Memory selection is automatic. The local validation run used --max-memory 48G on a 128 GiB machine.

Reasoning is on by default. For ordinary non-thinking responses, send "reasoning_effort": "none" through the OpenAI-compatible /v1/chat/completions API. Keep the service localhost-only unless you deliberately configure authentication and network access.

Pi

Merge a provider into your existing Pi models.json; do not replace your other providers. Use tcstacks/Qwen3.6-35B-A3B-Heretic-Splash as the model ID and http://127.0.0.1:8000/v1 as the base URL.

The important compatibility settings are:

{
  "api": "openai-completions",
  "apiKey": "local-placeholder",
  "compat": {
    "supportsStore": false,
    "supportsDeveloperRole": false,
    "supportsReasoningEffort": true,
    "supportsStrictMode": false,
    "maxTokensField": "max_tokens"
  }
}

For the model entry, set reasoning to true, thinkingLevelMap to {"off":"none"}, input to ["text"], contextWindow to 65536, and maxTokens to 8192. Select that provider/model in Pi. The compatibility settings and thinking-off mapping were exercised with actual Pi tool calls.

What was verified

  • All 625 sections across the 42 stock target files were reproduced byte-for-byte before converting Heretic.
  • All 1,757 language-model tensors were consumed by the conversion.
  • Quantized projections passed an independent inverse packing check for integer codes and BF16 scales/biases.
  • All 625 emitted target sections were read back and hashed.
  • The installed Splash validator verified all 56 artifact sizes and SHA-256 hashes.
  • A tool-call/tool-result API round trip passed.
  • Pi completed a real read/edit/execute task; an independent checker passed five interval-merging cases, including nested/touching intervals and preservation of caller input.
  • Thinking-off requests emitted zero reasoning tokens.

The 333 vision tensors in the source checkpoint were identical to the stock reference, so the stock vision artifact is preserved. Image and video inputs were not runtime-tested for this release. No broad capability, perplexity, or refusal-rate evaluation was performed. Upstream quality/refusal claims are not claims measured here.

Performance observations, not an isolated comparison

Observed on Apple M5 Max (40 GPU cores, 128 GiB), macOS 27, Splash 1.0:

Workload Samples Median decode rate
Short prompts, thinking off 3 275.5 tokens/s
Short prompts, thinking on 8 225.0 tokens/s
Approximately 35K-token prompts, thinking on 4 184.7 tokens/s

Thinking-on requests had a 1,024-token output cap. Thinking-off outputs stopped at model-selected lengths. All cold requests reported zero cached prompt tokens. These are generation rates, excluding prompt processing, not end-to-end throughput.

The retained draft accepted 21,269 of 38,262 proposed tokens (55.6%) across one warmup plus 31 scored requests. There were no reported Metal or capacity failures during that Heretic run.

A separate stock service restarted during the paired measurement, and an attempted repeat detected four unrelated stock requests. These observations do not establish an isolated speedup or slowdown against stock Splash and are not a hardware performance guarantee.

Reproduce the conversion

The converter intentionally targets this compatible Qwen3.6 MoE architecture and packed layout; it is not a universal model exporter. It requires the stock MLX source, Heretic MLX source, and stock packed package. Rebuilding therefore needs substantially more disk space than downloading this finished package.

Use Python 3.11 or newer. Download this repository's adapt_heretic.py and requirements-conversion.txt, then:

python3 -m venv .venv
. .venv/bin/activate
python -m pip install -r requirements-conversion.txt

hf download mlx-community/Qwen3.6-35B-A3B-4bit \
  --revision 38740b847e4cb78f352aba30aa41c76e08e6eb46 --local-dir stock-source
hf download froggeric/Qwen3.6-35B-A3B-Uncensored-Heretic-MLX-4bit \
  --revision 32939a6cef2750f18ccf352443f22f2e4dfe3613 --local-dir heretic-source
hf download incoai/Qwen3.6-35B-A3B-Splash \
  --revision 0f4714b2db37b5f3c42a10de07281e74f88e4adc --local-dir stock-splash

python adapt_heretic.py --verify-reference \
  --stock-source stock-source --source heretic-source \
  --template stock-splash --output local-stock-proof.json

python adapt_heretic.py --convert \
  --stock-source stock-source --source heretic-source \
  --template stock-splash --verification local-stock-proof.json \
  --output rebuilt-heretic-splash

Use hf download --local-dir as shown: the converter reads the pinned revision from Hugging Face's local download metadata. The output directory must not already exist. The published stock-packing-proof.json is sanitized historical evidence; generate your own local proof for conversion as above.

The converter emits runtime artifacts and provenance. If redistributing a rebuild, also preserve the applicable licenses, attribution/modification notices, and release documentation; the conversion command does not generate those legal/documentation files.

Files and provenance

  • target/: repacked Heretic target, 42 native files.
  • draft/: unchanged stock DFlash2 draft, 7 files.
  • vision/: preserved stock vision encoder.
  • tokenizer/: compatible vocabulary and Heretic tokenizer/template metadata.
  • manifest.json: source revisions, packed format, and hashes of 56 runtime artifacts.
  • layout.json: native section layout.
  • conversion-audit.json: per-section hashes and conversion provenance.
  • stock-packing-proof.json: byte-for-byte stock reference verification.
  • adapt_heretic.py: exact converter used for the target artifacts.

License and credits

Apache-2.0; see LICENSE and NOTICE.

The stock Splash package declares Apache-2.0 for every bundled component. No Splash runtime binary, private configuration, authentication token, or private coding transcript is included here.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tcstacks/Qwen3.6-35B-A3B-Heretic-Splash