Qwen3.6-14B-A3B-FableVibes-mlx-q6

MLX 6-bit quantization of tvall43/Qwen3.6-14B-A3B-FableVibes, for local inference on Apple Silicon.

Credit / original model

This repo is only a quantized MLX conversion. All credit for the model itself goes to the original author, tvall43. Please see and cite the original model card.

The base is a REAP-pruned Qwen3.6-35B-A3B reduced to ~14B total / ~3B active (90 experts, 8 active), recovered with a QLoRA distill of Claude Fable 5 reasoning traces. It uses the Qwen3.5 hybrid architecture (GatedDeltaNet linear attention + full attention + MoE) and emits <think>...</think> reasoning.

What this conversion did

  • Fused the routed-MoE experts from per-expert tensors (experts.{i}.{gate,up,down}_proj) into mlx-lm's stacked experts.gate_up_proj / experts.down_proj format.
  • Quantized to 6-bit, group size 64 with mlx-lm.
  • ~10 GB; runs on a 16 GB Apple Silicon Mac.

Usage

uv run python -m mlx_lm generate \
  --model khanh2023/Qwen3.6-14B-A3B-FableVibes-mlx-q6 \
  --prompt "Solve: ..."

Notes

  • MoE sparsity (3B active/token) makes decode fast (46 tok/s on an M4) despite 14B total params.
  • 6-bit preserves more exactness than q4 on strict reasoning, at ~10 GB (needs a raised Metal wired limit on 16 GB). A smaller -mlx-q4 variant is also available.
Downloads last month
466
Safetensors
Model size
14B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for khanh2023/Qwen3.6-14B-A3B-FableVibes-mlx-q6

Quantized
(3)
this model

Collection including khanh2023/Qwen3.6-14B-A3B-FableVibes-mlx-q6