RavenX-CyberAgent-Qwen3.6-35B-A3B-Splash-Q4

RavenX-CyberAgent-Qwen3.6-35B-A3B converted to the Splash schema-4 splash-packed-q4-moe format for Apple Silicon. This repository is a Splash runtime package, not an MLX checkpoint, GGUF file, or Transformers checkpoint.

The target and tokenizer come from the Ck-TnT MLX 4-bit RavenX checkpoint, which derives from the original RavenX model. The compatible DFlash 2 draft and vision weights are reused from Inco AI's Qwen3.6 Splash package. All 40 converted language-model layer files were built from RavenX's weights. The reused official Splash weight assets are the draft model and vision encoder.

Run with Splash

Install Splash on an Apple Silicon Mac, then replace YOUR_USERNAME with this repository's owner:

brew install incoai/tap/splash
splash serve --model YOUR_USERNAME/RavenX-CyberAgent-Qwen3.6-35B-A3B-Splash-Q4 --max-context 8K

After Splash prints Ready, open http://127.0.0.1:8000 or send a chat completion request. The 8K value above is a server limit, not a claim that an 8K-token prompt has been tested.

Conversion and package

The MLX affine 4-bit codes, BF16 scales, and BF16 biases were repacked into Splash's Q4 tile layout without a second quantization. Expert routers and the shared-expert scalar gate remain Q8. The architecture has 40 layers in a repeating three-GDN/one-full-attention pattern, 256 routed experts, and eight active experts per token. The package uses RavenX's own tokenizer and chat template.

Component Size Source
target/ (42 files) 19,541,082,112 bytes (18.199 GiB) RavenX MLX 4-bit checkpoint
Full package 20,947,991,647 bytes (19.509 GiB) before this card and LICENSE Target, RavenX tokenizer, compatible draft and vision assets

manifest.json lists all 56 runtime artifacts with their sizes and SHA-256 hashes. layout.json describes the target sections. The converter validated all required files, headers, layer sizes, section offsets, and artifact hashes against the official Qwen3.6 Splash format. Splash's installer manifest and artifact checks passed. Forty-three sampled decoded Q4/Q8 rows matched the dequantized MLX source exactly; this does not measure error against the original BF16 checkpoint.

Runtime validation and limits

A user-reported local Splash test on an M5 Pro Mac with 48 GB unified memory loaded this package and generated a short text response with an 8K configured context limit. The response and memory measurements were not recorded. Long prompts, RavenX-specific tool calling, and image quality have not been evaluated. The reused draft is shape-compatible, but its acceptance rate and speed with this fine-tune have not been measured. Current Splash requires a draft model; draft-free generation has not been tested.

License and credits

Apache-2.0. The RavenX source, its MLX 4-bit conversion, and Inco AI's Splash reference package each list Apache-2.0. See their linked model cards for their authors, provenance, and intended use. The Splash runtime is by Inco AI.

This conversion used a Qwen3.6 MoE extension to publicExcess/splash-converter; link the implementation pull request here when it is published.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yannipang/RavenX-CyberAgent-Qwen3.6-35B-A3B-Splash-Q4