Qwen3.8-27B-Uncensored-MXFP4-awq

MXFP4 quantization of orcarouter/Qwen3.8-27B-Uncensored, made with AMD Quark and exported as an AWQ-style checkpoint for serving with vLLM on AMD RDNA4 GPUs (gfx1201 / RX 9000 series) with ROCm.

What I used to run it

Start with this image: https://hub.docker.com/r/stilldeadcode/vllm-radiance

Add this if needed (needed at the time of testing): https://codeberg.org/ggz14/radiance-vllm-mxfp4

NOTE: fp8_mtp.py was already ran on this model, it is not needed to run again.

Ran Livebench Math for quality check (GLM 5.3 was grader): Final scoreboard (official 182q math, xhigh, 150k tokens) AMPS_Hard 77.0% math_comp 97.7% olympiad 100% Overall 86.5%

Model Details

Source model orcarouter/Qwen3.8-27B-Uncensored
Architecture Qwen3.5-27B (Qwen3_5ForConditionalGeneration)
Quantization tool AMD Quark
Quantization MXFP4 (Quark export, quant_method: quark)
KV cache Post-RoPE KV quantization enabled
Context length 262,144
Hidden size / layers 5120 / 64
Attention 24 heads, 4 KV heads (GQA)
Vocab size 248,320
Multimodal Vision tower present (excluded from quantization)

Quantization excludes lm_head and the vision tower, so those run at higher precision.

Usage

Serve with vLLM:

vllm serve just1moremodel/Qwen3.8-27B-Uncensored-MXFP4-awq \
  --port 8081

Requires a vLLM build with Quark/MXFP4 support (ROCm on RDNA4 recommended).

Intended Use

  • Uncensored/abliterated fine-tune — intended for research, creative writing, and local inference where refusal behavior is not desired.
  • Single-file model.safetensors (~19 GB), sized to fit consumer GPUs with 24 GB+ VRAM (depending on context length and offloading).

Limitations

  • MXFP4 is an aggressive quantization; expect some quality loss vs bf16/fp8.
  • Uncensored models may produce harmful or objectionable output. You are responsible for how you use this model.
  • Not tested for production or safety-critical use.

License

Apache-2.0, inherited from the base model.

Downloads last month
920
Safetensors
Model size
16B params
Tensor type
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for just1moremodel/Qwen3.8-27B-Uncensored-MXFP4-awq

Base model

Qwen/Qwen3.8-27B
Quantized
(57)
this model