Qwen3.6-27B-AEON-Ultimate-Uncensored — Q6_K GGUF + MTP

GGUF Q6_K quantization of AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16 with MTP (Multi-Token Prediction) weights included.

Key Features

  • Q6_K quantization — near-lossless quality, practical for consumer hardware
  • MTP weights included (866 tensors) — enables speculative decoding via --spec-type draft-mtp in llama.cpp for significantly faster inference
  • Vision supported — mmproj file (931MB) available in separate repo for multimodal image input

Model Details

Property Value
Base Model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
Architecture Qwen 3.6 (Hybrid Linear Attention + Full Attention, 64 layers)
Parameters 27B
Quantization Q6_K
Main Model Size 21.4 GB
mmproj Size 931 MB
Main Model Tensors 866 (including MTP extension)
mmproj Tensors 334 (27 ViT blocks + merger + patch embed + position embed)
Context Length 262,144 tokens (training)
License Apache 2.0

Quantization Process

  1. BF16 original → FP16 GGUF conversion (V100 does not support BF16)
  2. FP16 GGUF → Q6_K quantization with MTP weights preserved

How to Use

llama.cpp (recommended)

# Text-only
llama-server \
  -m qwen3.6-27b-Q6K-mtp.gguf \
  -ngl 99 \
  -c 80000 \
  --host 0.0.0.0 \
  --port 8080 \
  --spec-type draft-mtp

# With vision (download mmproj from https://huggingface.co/pyys/Qwen3.6-27B-mmproj-GGUF)
llama-server \
  -m qwen3.6-27b-Q6K-mtp.gguf \
  --mmproj qwen3.6-27b-mmproj.gguf \
  -ngl 99 \
  -c 80000 \
  --host 0.0.0.0 \
  --port 8080 \
  --spec-type draft-mtp

MTP (Multi-Token Prediction)

MTP enables speculative decoding without a separate draft model. The MTP weights are baked into this GGUF file. Use --spec-type draft-mtp to activate.

Performance

Tested on V100 SXM2 32GB:

Configuration TPS Notes
Q6_K + MTP ~31.5 t/s MTP acceptance rate ~41%
Q6_K + MTP (via OpenAI-compatible API) ~39.5 t/s Acceptance rate ~61%

Files

File Size Description
qwen3.6-27b-Q6K-mtp.gguf 21.4 GB Main model (Q6_K + MTP weights)
qwen3.6-27b-mmproj.gguf 931 MB Vision projector (separate repo)

Notes

  • V100 GPUs do not support BF16 — this model was converted via FP16 intermediate
  • For single-GPU deployment, use --split-mode none -mg 0 to keep all weights on one GPU
  • KV cache of 80K tokens fits within a single V100 32GB alongside the model

Credits

  • Original model: AEON-7 — Qwen3.6-27B-AEON-Ultimate-Uncensored
  • Base architecture: Qwen — Qwen3.6-27B
  • Quantization tooling: llama.cpp
Downloads last month
448
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pyys/Qwen3.6-27B-AEON-Ultimate-Uncensored-Q6K-MTP-GGUF

Base model

Qwen/Qwen3.6-27B
Quantized
(41)
this model