RavenX-CyberAgent-Qwen3.6-35B-A3B-Splash-Q4
RavenX-CyberAgent-Qwen3.6-35B-A3B converted to the Splash schema-4
splash-packed-q4-moe format for Apple Silicon. This repository is a Splash
runtime package, not an MLX checkpoint, GGUF file, or Transformers checkpoint.
The target and tokenizer come from the Ck-TnT MLX 4-bit RavenX checkpoint, which derives from the original RavenX model. The compatible DFlash 2 draft and vision weights are reused from Inco AI's Qwen3.6 Splash package. All 40 converted language-model layer files were built from RavenX's weights. The reused official Splash weight assets are the draft model and vision encoder.
Run with Splash
Install Splash on an Apple Silicon Mac,
then replace YOUR_USERNAME with this repository's owner:
brew install incoai/tap/splash
splash serve --model YOUR_USERNAME/RavenX-CyberAgent-Qwen3.6-35B-A3B-Splash-Q4 --max-context 8K
After Splash prints Ready, open http://127.0.0.1:8000 or send a chat
completion request. The 8K value above is a server limit, not a claim that
an 8K-token prompt has been tested.
Conversion and package
The MLX affine 4-bit codes, BF16 scales, and BF16 biases were repacked into Splash's Q4 tile layout without a second quantization. Expert routers and the shared-expert scalar gate remain Q8. The architecture has 40 layers in a repeating three-GDN/one-full-attention pattern, 256 routed experts, and eight active experts per token. The package uses RavenX's own tokenizer and chat template.
| Component | Size | Source |
|---|---|---|
target/ (42 files) |
19,541,082,112 bytes (18.199 GiB) | RavenX MLX 4-bit checkpoint |
| Full package | 20,947,991,647 bytes (19.509 GiB) before this card and LICENSE |
Target, RavenX tokenizer, compatible draft and vision assets |
manifest.json lists all 56 runtime artifacts with their sizes and SHA-256
hashes. layout.json describes the target sections. The converter validated
all required files, headers, layer sizes, section offsets, and artifact
hashes against the official Qwen3.6 Splash format. Splash's installer manifest
and artifact checks passed. Forty-three sampled decoded Q4/Q8 rows matched
the dequantized MLX source exactly; this does not measure error against the
original BF16 checkpoint.
Runtime validation and limits
A user-reported local Splash test on an M5 Pro Mac with 48 GB unified memory loaded this package and generated a short text response with an 8K configured context limit. The response and memory measurements were not recorded. Long prompts, RavenX-specific tool calling, and image quality have not been evaluated. The reused draft is shape-compatible, but its acceptance rate and speed with this fine-tune have not been measured. Current Splash requires a draft model; draft-free generation has not been tested.
License and credits
Apache-2.0. The RavenX source, its MLX 4-bit conversion, and Inco AI's Splash reference package each list Apache-2.0. See their linked model cards for their authors, provenance, and intended use. The Splash runtime is by Inco AI.
This conversion used a Qwen3.6 MoE extension to publicExcess/splash-converter; link the implementation pull request here when it is published.