khanh2023's picture
MLX q6 conversion of tvall43/Qwen3.6-14B-A3B-FableVibes
cd8f254 verified
|
Raw History Blame Contribute Delete
1.75 kB
---
license: apache-2.0
base_model: tvall43/Qwen3.6-14B-A3B-FableVibes
pipeline_tag: text-generation
library_name: mlx
tags:
- mlx
- apple-silicon
- moe
- qwen3.5
- reasoning
---
# Qwen3.6-14B-A3B-FableVibes-mlx-q6
**MLX 6-bit** quantization of [**tvall43/Qwen3.6-14B-A3B-FableVibes**](https://huggingface.co/tvall43/Qwen3.6-14B-A3B-FableVibes), for local inference on Apple Silicon.
## Credit / original model
This repo is **only a quantized MLX conversion**. All credit for the model itself goes to the original author, **[tvall43](https://huggingface.co/tvall43)**. Please see and cite the [original model card](https://huggingface.co/tvall43/Qwen3.6-14B-A3B-FableVibes).
The base is a **REAP-pruned** [Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) reduced to ~14B total / ~3B active (90 experts, 8 active), recovered with a QLoRA distill of Claude Fable 5 reasoning traces. It uses the Qwen3.5 hybrid architecture (GatedDeltaNet linear attention + full attention + MoE) and emits `<think>...</think>` reasoning.
## What this conversion did
- **Fused the routed-MoE experts** from per-expert tensors (`experts.{i}.{gate,up,down}_proj`) into mlx-lm's stacked `experts.gate_up_proj` / `experts.down_proj` format.
- Quantized to **6-bit, group size 64** with `mlx-lm`.
- ~10 GB; runs on a 16 GB Apple Silicon Mac.
## Usage
```bash
uv run python -m mlx_lm generate \
--model khanh2023/Qwen3.6-14B-A3B-FableVibes-mlx-q6 \
--prompt "Solve: ..."
```
## Notes
- MoE sparsity (~3B active/token) makes decode fast (~46 tok/s on an M4) despite 14B total params.
- 6-bit preserves more exactness than q4 on strict reasoning, at ~10 GB (needs a raised Metal wired limit on 16 GB). A smaller `-mlx-q4` variant is also available.