mlx-community-Qwythos-9B-v2-OptiQ-4bit-MTPLX

MTPLX-branded multi-token-prediction model optimized for Apple Silicon (MLX).
Forged with MTPLX Forge from mlx-community/Qwythos-9B-v2-OptiQ-4bit.


Verification & Benchmarks

Metric Result
Speedup 1.85× vs. autoregressive baseline
Throughput 67.1 tok/s (Baseline: 36.3 tok/s)
Best Depth D2
Mean Acceptance Rate 93% at D2
Hardware / OS Apple M4 Pro · macOS 26.6.2
Sampler Config temp=0.6, top_p=0.95, top_k=20

Refer to mtplx_runtime.json for the complete verification log.


Usage

MTPLX detects and initializes the model automatically once downloaded:

# Pull model weights
mtplx pull nRanzo/mlx-community-Qwythos-9B-v2-OptiQ-4bit-MTPLX

# Launch interactive session
mtplx start chat

Project & Implementation

For benchmarks, speculative decoding pipelines, and testing methodology, visit the MLX-MTP Repository.


Attribution & License

Downloads last month
140
Safetensors
Model size
9B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nRanzo/mlx-community-Qwythos-9B-v2-OptiQ-4bit-MTPLX