Qwen3.6-35B-A3B Heretic — MLX BF16, Vision + Native MTP

Full multimodal MLX-BF16 conversion of llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved, pinned to revision 9599ac17d26daf33daf0fdd8f6c897ff4c6dc89a.

This build preserves all three required components:

  • the 40-layer Qwen3.6-35B-A3B MoE language backbone;
  • the complete 333-tensor vision tower and image/video processor metadata;
  • the embedded one-layer Native-MTP head (no external draft model).

The source has 1,045 tensors: 333 vision tensors, 19 Native-MTP tensors, and 693 remaining language tensors. The MLX conversion has 1,086 tensors because the 40 backbone MoE gate/up tensors and the one MTP MoE gate/up tensor are split into MLX runtime projections. All saved tensors remain BF16; no quantization is applied.

oMLX compatibility

Verified against oMLX commit 4cb5516d3de3184209cdfaa53369c8c33b6a91ba and its pinned mlx-vlm commit 78b96eb5462141447b9a6b4943ef553891da56dd. The verification requires:

  • strict oMLX VLM loading;
  • a bound language_model.mtp head and mtp_forward runtime method;
  • actual image preprocessing and vision-embedding generation;
  • a short end-to-end image-conditioned generation.

Copy or clone the repository into the oMLX model directory, select it as a Vision-Language Model, and enable Native MTP in the model settings. Do not select an external draft model for Native MTP.

BF16 is not eligible for oMLX Qwen ANE Prompt Processing, which currently accepts affine 4/5/6/8-bit projections with group size 64 or 128. This build therefore uses the GPU for BF16 prompt processing.

Reproducibility

The conversion project records its source pins and includes a static verifier plus a runtime smoke test. The generated conversion-manifest.json records the exact source revision and tensor counts.

Safety and licensing

This is an uncensored/heretic derivative. Outputs may be inaccurate, offensive, unsafe, or unlawful for a particular use. Apply independent safety controls and human review. Licensing and use restrictions of the upstream checkpoint and all of its base models continue to apply; consult the linked source repository before redistribution or deployment.

Downloads last month
334
Safetensors
Model size
36B params
Tensor type
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for trigger2k25/Qwen3.6-35B-A3B-heretic-BF16-mtp