Qwen3.8-27B-TURBO-Heretic-NM-DAU (MLX 8-bit)

MLX conversion of DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU (a heretic finetune of Qwen/Qwen3.8-27B). The language model was converted text-only with mlx-lm, the source checkpoint's BF16 vision tower was grafted back in as vision_tower.*, the result validated for the Splash engine (family Qwen3.8-27B), and the chat template swapped to Sharp v22.5.0.

Quantization

  • Language model linears (incl. lm_head): affine 8-bit, group size 64
  • Token embeddings, GDN A_log/dt_bias, norms: BF16
  • Vision tower (vision_tower.*, 333 tensors): BF16, unquantized — grafted from the source checkpoint's model.visual.* weights
  • MTP weights from the source checkpoint were dropped

Bits per weight: 8.501. Total size: ~28 GB.

Layout

7 safetensors shards: shard 1 holds the vision tower, shards 2–7 the language model, matching the layout of mlx-community/Qwen3.8-27B-4bit.

Build

  1. mlx_lm convert -q --q-bits 8 (text-only output; mlx-lm 0.32.0 drops the model.visual tower for qwen3_5)
  2. Tower grafted byte-for-byte from the source BF16 shards as vision_tower.*, vision_config and processor files restored from the source repo

Chat template

Sharp v22.5.0 (qwen3.8-froggeric-v22.5.0) replaces the stock Qwen VL template: terseness block appended after your system prompt, thinking retention for prefix-cache hits, tool-call token parity. Per-request knobs via chat_template_kwargs: {"terse": false}, reasoning_effort, enable_thinking.

Splash

Passes the Splash engine model-check as family Qwen3.8-27B:

splash serve --model blackfan23/Qwen3.8-27B-TURBO-Heretic-NM-DAU-MLX-8bit
Downloads last month
112
Safetensors
Model size
27B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for blackfan23/Qwen3.8-27B-TURBO-Heretic-NM-DAU-MLX-8bit