Qwen3.8-27B Ternary PQ2_0 (2.13 bpw) with Native MTP

2-bit / 2.13 bpw ternary quantization of Qwen3.8-27B (Prism ML Ternary Bonsai 2 architecture) with native 1-layer MTP speculative drafting.

  • Format: PQ2_0 (group 128, 34 bytes/block, ~2.13 bpw)
  • Model Size: 7.66 GB
  • MTP Drafter: Native 1-layer MTP head (Q8_0 projection matrices fine-tuned on Bonsai 2)
  • Target Hardware: Single 24 GB GPU (RTX 3090 / 3090 Ti) with 245K+ context capacity under full offload
  • Runtime: Optimized for llamAmpere (branch feature/v0.3.1-bonsai2)

Usage with llamAmpere

./build-sm86/bin/llama-server \
  -m Qwen3.8-27B-Ternary-PQ2_0-MTP.gguf \
  -c 32768 -b 4096 -ub 1024 -t 8 -ngl 99 -fa on -ctk q8_0 -ctv turbo3 \
  --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0
Downloads last month
1,076
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including jakeatx/Qwen3.8-27B-Ternary-PQ2_0-MTP-GGUF