Download README.md from plunderstruck/Qwen3.6-27B-OBLITERATED-MTP-ROCmFP4-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 4.28 kB
-
https://huggingface.co/plunderstruck/Qwen3.6-27B-OBLITERATED-MTP-ROCmFP4-GGUF/resolve/7f01c52cf0e749dad8994b48638c6fd8a4d2432d/README.md
- Command line
-
hf download hf://plunderstruck/Qwen3.6-27B-OBLITERATED-MTP-ROCmFP4-GGUF@7f01c52cf0e749dad8994b48638c6fd8a4d2432d/README.md
-
curl -L -o README.md https://huggingface.co/plunderstruck/Qwen3.6-27B-OBLITERATED-MTP-ROCmFP4-GGUF/resolve/7f01c52cf0e749dad8994b48638c6fd8a4d2432d/README.md
base_model: OBLITERATUS/Qwen3.6-27B-OBLITERATED
base_model_relation: quantized
license: apache-2.0
library_name: gguf
tags:
- gguf
- rocmfp4
- qwen3.6
- obliterated
- abliterated
- uncensored
- 27b
- mtp
- speculative-decoding
- strix-halo
- amd
- rocm
- vulkan
language:
- en
Qwen3.6-27B-OBLITERATED-MTP — ROCmFP4 STRIX
🚧 Work in progress — files not uploaded yet
This card is up so you know it's coming. The GGUFs aren't here yet — the BF16 source is downloading and the quant/measurement pipeline is queued behind some other uploads. Check back; this notice gets replaced with the files + measured numbers when it's done.
Experimental AMD Strix Halo (gfx1151) quant of OBLITERATUS/Qwen3.6-27B-OBLITERATED — an abliterated / uncensored Qwen3.6-27B (refusal direction removed, with source-weight interpolation to retain capability) — in the custom ROCmFP4 4-bit format, with an MTP / next-token-prediction head grafted in for self-speculative decoding on a single APU.
⚠️ Ignore HuggingFace's auto-detected quant badge ("F16"/16-bit) — it's wrong. HF can't read the custom ROCmFP4 tensor types and mislabels the file by its f16 embeddings. These will be ~4.5 bpw 4-bit ROCmFP4 files, not 16-bit. Pick by filename in Files and versions.
Requires the ROCmFP4 fork (public) — not stock llama.cpp
Uses the ROCmFP4 tensor types (
q4_0_rocmfp4,q4_0_rocmfp4_fast). Stock llama.cpp, LM Studio, Ollama, etc. cannot load it. Build/run withcharlie12345/rocmfp4-llama(mtp-rocmfp4-strix).
Lineage
this ROCmFP4 quant ──quantized──▶ OBLITERATUS/Qwen3.6-27B-OBLITERATED ──abliterated──▶ Qwen/Qwen3.6-27B
It's a 4-bit Strix-Halo quant of OBLITERATUS's abliterated 27B. The abliteration (refusal-direction removal + source interpolation) is upstream work; we only do the ROCmFP4 quant + MTP graft.
Planned recipe (what these will be)
- Genuine f16 token embeddings — quantized from the BF16 safetensors, so the f16 embeddings are real f16 (not Q8→f16 "fake-f16").
- Grafted MTP head — a
nextnhead transplanted from a Qwen3.6-27B-MTP BF16 donor (output-lossless; it only affects draft speed), so it runs self-speculative on the fork. - Two output-head variants — a base (4-bit head, fastest) and a
Q6_Khead (a notch more faithful). - imatrix: measured, then decided. This is a dense Qwen3.6-27B — the same architecture class as our Qwopus-Coder, the one model where a code-weighted imatrix worsened perplexity. So we'll build both and pick by measured KL + perplexity vs the BF16, rather than assume. Whatever ships, the card will say which won and show the numbers.
Status
🚧 Building. No files yet, and no quality measurement yet — nothing here is a quality claim. When it lands it'll carry the same honest scope as our other cards: a KL/PPL fidelity comparison vs BF16, not an absolute benchmark.
Sibling ROCmFP4 Strix Halo models
- Qwen3.6-27B-MTP · Qwen3.6-35B-A3B-MTP · Qwopus3.6-27B-Coder-MTP
- Qwen3.6-40B-Deckard-MTP · Qwen3-Coder-Next · Nex-N2-mini
Credits & license
- Base model:
OBLITERATUS/Qwen3.6-27B-OBLITERATED(Apache-2.0), an abliterated derivative ofQwen/Qwen3.6-27B(Qwen team). A derivative quantization — verify the base terms before redistribution/use. - ROCmFP4 format & runtime:
charlie12345/rocmfp4-llama(based on llama.cpp, MIT).