File size: 4,280 Bytes
4867049
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
---
base_model: OBLITERATUS/Qwen3.6-27B-OBLITERATED
base_model_relation: quantized
license: apache-2.0
library_name: gguf
tags:
- gguf
- rocmfp4
- qwen3.6
- obliterated
- abliterated
- uncensored
- 27b
- mtp
- speculative-decoding
- strix-halo
- amd
- rocm
- vulkan
language:
- en
---

# Qwen3.6-27B-OBLITERATED-MTP — ROCmFP4 STRIX

> ## 🚧 Work in progress — files not uploaded yet
> This card is up so you know it's coming. The **GGUFs aren't here yet** — the BF16 source is downloading
> and the quant/measurement pipeline is queued behind some other uploads. Check back; this notice gets
> replaced with the files + measured numbers when it's done.

Experimental **AMD Strix Halo (gfx1151)** quant of [**OBLITERATUS/Qwen3.6-27B-OBLITERATED**](https://huggingface.co/OBLITERATUS/Qwen3.6-27B-OBLITERATED)
— an **abliterated / uncensored** Qwen3.6-27B (refusal direction removed, with source-weight interpolation
to retain capability) — in the custom **ROCmFP4** 4-bit format, with an MTP / next-token-prediction head
**grafted in** for self-speculative decoding on a single APU.

> **⚠️ Ignore HuggingFace's auto-detected quant badge ("F16"/16-bit) — it's wrong.**
> HF can't read the custom ROCmFP4 tensor types and mislabels the file by its f16 embeddings. **These will
> be ~4.5 bpw 4-bit ROCmFP4 files, not 16-bit.** Pick by filename in *Files and versions*.

> ## Requires the ROCmFP4 fork (public) — not stock llama.cpp
> Uses the **ROCmFP4** tensor types (`q4_0_rocmfp4`, `q4_0_rocmfp4_fast`). **Stock llama.cpp, LM Studio,
> Ollama, etc. cannot load it.** Build/run with
> **[`charlie12345/rocmfp4-llama`](https://github.com/charlie12345/rocmfp4-llama)** (`mtp-rocmfp4-strix`).

## Lineage

```
this ROCmFP4 quant  ──quantized──▶  OBLITERATUS/Qwen3.6-27B-OBLITERATED  ──abliterated──▶  Qwen/Qwen3.6-27B
```

It's a 4-bit Strix-Halo quant of OBLITERATUS's abliterated 27B. The abliteration (refusal-direction
removal + source interpolation) is upstream work; we only do the ROCmFP4 quant + MTP graft.

## Planned recipe (what these will be)

- **Genuine f16 token embeddings** — quantized **from the BF16 safetensors**, so the f16 embeddings are
  real f16 (not Q8→f16 "fake-f16").
- **Grafted MTP head** — a `nextn` head transplanted from a Qwen3.6-27B-MTP BF16 donor (output-lossless;
  it only affects draft *speed*), so it runs self-speculative on the fork.
- **Two output-head variants** — a base (4-bit head, fastest) and a **`Q6_K` head** (a notch more faithful).
- **imatrix: measured, then decided.** This is a *dense* Qwen3.6-27B — the same architecture class as our
  [Qwopus-Coder](https://huggingface.co/plunderstruck/Qwopus3.6-27B-Coder-MTP-ROCmFP4-GGUF), the one model
  where a code-weighted imatrix *worsened* perplexity. So we'll build both and pick by **measured KL +
  perplexity vs the BF16**, rather than assume. Whatever ships, the card will say which won and show the numbers.

## Status

🚧 Building. No files yet, and **no quality measurement yet** — nothing here is a quality claim. When it
lands it'll carry the same honest scope as our other cards: a KL/PPL fidelity comparison vs BF16, not an
absolute benchmark.

## Sibling ROCmFP4 Strix Halo models

- [Qwen3.6-27B-MTP](https://huggingface.co/plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF) · [Qwen3.6-35B-A3B-MTP](https://huggingface.co/plunderstruck/Qwen3.6-35B-A3B-MTP-ROCmFP4-GGUF) · [Qwopus3.6-27B-Coder-MTP](https://huggingface.co/plunderstruck/Qwopus3.6-27B-Coder-MTP-ROCmFP4-GGUF)
- [Qwen3.6-40B-Deckard-MTP](https://huggingface.co/plunderstruck/Qwen3.6-40B-Deckard-MTP-ROCmFP4-GGUF) · [Qwen3-Coder-Next](https://huggingface.co/plunderstruck/Qwen3-Coder-Next-ROCmFP4-GGUF) · [Nex-N2-mini](https://huggingface.co/plunderstruck/Nex-N2-mini-ROCmFP4-GGUF)

## Credits & license

- **Base model:** [`OBLITERATUS/Qwen3.6-27B-OBLITERATED`](https://huggingface.co/OBLITERATUS/Qwen3.6-27B-OBLITERATED) (Apache-2.0), an abliterated derivative of [`Qwen/Qwen3.6-27B`](https://huggingface.co/Qwen/Qwen3.6-27B) (Qwen team). A derivative quantization — verify the base terms before redistribution/use.
- **ROCmFP4 format & runtime:** [`charlie12345/rocmfp4-llama`](https://github.com/charlie12345/rocmfp4-llama) (based on llama.cpp, MIT).