|
Download README.md from plunderstruck/Qwen3.6-27B-OBLITERATED-MTP-ROCmFP4-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 30.3 kB
-
https://huggingface.co/plunderstruck/Qwen3.6-27B-OBLITERATED-MTP-ROCmFP4-GGUF/resolve/3bd9cd7797ae679bee58df61770e405747998310/README.md
- Command line
-
hf download hf://plunderstruck/Qwen3.6-27B-OBLITERATED-MTP-ROCmFP4-GGUF@3bd9cd7797ae679bee58df61770e405747998310/README.md
-
curl -L -o README.md https://huggingface.co/plunderstruck/Qwen3.6-27B-OBLITERATED-MTP-ROCmFP4-GGUF/resolve/3bd9cd7797ae679bee58df61770e405747998310/README.md
30.3 kB
| base_model: OBLITERATUS/Qwen3.6-27B-OBLITERATED | |
| base_model_relation: quantized | |
| license: apache-2.0 | |
| library_name: gguf | |
| tags: | |
| - gguf | |
| - rocmfp4 | |
| - qwen3.6 | |
| - obliterated | |
| - abliterated | |
| - uncensored | |
| - 27b | |
| - mtp | |
| - speculative-decoding | |
| - strix-halo | |
| - amd | |
| - rocm | |
| - vulkan | |
| language: | |
| - en | |
| <div style="border:2px solid currentColor; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;"> | |
| <div style="border-bottom:1px solid currentColor; padding:6px 12px; font-size:11px; letter-spacing:3px; text-transform:uppercase; opacity:0.7; text-align:center;">PLUNDERSTRUCK // ROCmFP4 QUANTIZED MODEL // STRIX HALO · gfx1151</div> | |
| <div style="padding:14px; display:flex; flex-wrap:wrap; align-items:center; justify-content:center; gap:18px;"> | |
| <pre style="margin:0; flex:0 0 auto; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,monospace; font-size:5px; line-height:1.1; letter-spacing:0;"> | |
| ▗▇▇▇▇▇▇▇▖ | |
| ▗█▘▝██████▖ | |
| ▗▛ ▝██████▆▆▆▆▆▆▆▆▆▆▅ | |
| ▟▛ ▗█████████████████▙▖ | |
| ▄▄▄▄▄▟▛ ▟████████████████████▖ | |
| ▗██▌ ▚▖ ▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔▔█▘ | |
| ▗████▖ ▜▖ ▗█▘ | |
| ▜█████▙ ▜▆▆▆▆▆▆▆▆▆▆▆▆▆▆▆▀▀▀▀▀▜▙ | |
| ▜█████▙ ▝████████████▛ ▜▙ | |
| ▜█████▙ ▝██████████▛ ▃ ▜▙ | |
| ▀█████▙▖ ▝████████▘ ▟█▙ ▀▙ | |
| ▝██████▖ ▝▜█████▘ ▟███▙▂▂▂▂▐█ | |
| ▟███████▖ ▜███▘ ▗███████████▛ | |
| ▟█████████▄ ▜▛ ▗███████████▀ | |
| ▝█████▀ ▗▛ ▗██████▀▀▀▀▀▘ | |
| ▜██▘ ▗▛ ▟█████▛▘ | |
| ▜█▇▇▇▇▇▇▇▇▇█▖ ▟█████▛ | |
| ▝█▖ ▟█████▛ | |
| ▝███████▀ | |
| </pre> | |
| <div style="flex:0 1 auto; max-width:100%; text-align:center;"> | |
| <div style="font-size:23px; font-weight:800; letter-spacing:1px;">QWEN3.6-27B-OBLITERATED-MTP</div> | |
| <div style="font-size:12.5px; letter-spacing:1px; opacity:0.8; margin-top:5px;"><span style="white-space:nowrap;">4-BIT ROCmFP4</span> · <span style="white-space:nowrap;">ABLITERATED / UNCENSORED</span> · <span style="white-space:nowrap;">GRAFTED MTP SELF-SPECULATIVE DECODE</span> · <span style="white-space:nowrap;">VISION-CAPABLE</span> · <span style="white-space:nowrap;">SINGLE AMD APU</span></div> | |
| </div> | |
| </div> | |
| <table style="display:table; table-layout:fixed; width:100%; margin:0; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;"> | |
| <tr> | |
| <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">FORMAT</div><div style="font-weight:700;">ROCmFP4 4-BIT</div></td> | |
| <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">PRECISION</div><div style="font-weight:700;">~4.8 BPW</div></td> | |
| <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">SIZE</div><div style="font-weight:700;">~15 GB</div></td> | |
| <td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">CONTEXT</div><div style="font-weight:700;">262 K</div></td> | |
| </tr> | |
| <tr> | |
| <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">DRAFT</div><div style="font-weight:700;">MTP n-max 5 (GRAFTED)</div></td> | |
| <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">VISION</div><div style="font-weight:700;">QWEN3-VL</div></td> | |
| <td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">BACKEND</div><div style="font-weight:700;">VULKAN0</div></td> | |
| <td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">LICENSE</div><div style="font-weight:700;">APACHE-2.0</div></td> | |
| </tr> | |
| </table> | |
| </div> | |
| <div style="border:2px solid #dc2626; padding:10px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px; margin:14px 0;"> | |
| <b style="color:#dc2626; letter-spacing:1px;">⚠ REQUIRES THE ROCmFP4 FORK</b><br> | |
| The custom <code>q4_0_rocmfp4</code> / <code>q4_0_rocmfp4_fast</code> tensor types <b>will not load in stock llama.cpp, LM Studio, or Ollama</b>. Build/run with <a href="https://github.com/charlie12345/ROCmFPX">charlie12345/ROCmFPX</a> · branch <code>mtp-rocmfp4-strix</code>. | |
| </div> | |
| <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:14px 0; opacity:0.85;"> | |
| <b>NOTE //</b> Ignore HuggingFace's auto-detected "F16"/16-bit badge — its parser can't read ROCmFP4 and mislabels by the f16 embeddings. These are <b>~4.8 bpw 4-bit</b> files; pick by filename. | |
| </div> | |
| <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:14px 0; opacity:0.85;"> | |
| <b>NOTE //</b> <b>Uncensored:</b> this is an <i>abliterated</i> model — the refusal direction was removed (with source-weight interpolation to retain capability) upstream by OBLITERATUS. It will answer prompts a stock Qwen3.6 would refuse. Abliteration is upstream work; verify behavior before relying on it. | |
| </div> | |
| <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">01</span> · FILES</div> | |
| <div style="overflow:hidden; border-radius:0;"> | |
| <table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;"> | |
| <thead><tr> | |
| <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">File</th> | |
| <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Size</th> | |
| <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Output head</th> | |
| <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Pick if</th> | |
| </tr></thead> | |
| <tbody> | |
| <tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-STRIX-embF16-imatrix-headQ6.gguf</code> ★</td><td style="border:1px solid currentColor; padding:7px 10px;">~15.5 GB</td><td style="border:1px solid currentColor; padding:7px 10px;">Q6_K</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>the one build</b> — best speed/quality balance: f16 embeddings + Q6 output head on the fast single-scale body</td></tr> | |
| </tbody> | |
| </table> | |
| </div> | |
| One file — the **best speed/quality balance** in ROCmFP4 for Strix Halo. It keeps the two quality levers that are actually *felt* — genuine **f16 token embeddings** (from BF16) and a **Q6_K output head** — on the fast single-scale `q4_0_rocmfp4_fast` body + a general+code-calibrated imatrix + the **grafted MTP draft head**. Repo also bundles the **[`mmproj-F32.gguf`](./mmproj-F32.gguf)** Qwen3-VL vision projector and **[`chat_template.jinja`](./chat_template.jinja)** (froggeric's unified Qwen3.6 template — tool calls + inline `<|think_off|>`/`<|think_on|>` + vision). | |
| <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">02</span> · QUICK START</div> | |
| Run from the folder holding the `.gguf` + `chat_template.jinja`: | |
| ```bash | |
| env HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \ | |
| llama-server \ | |
| -m Qwen3.6-27B-OBLITERATED-MTP-ROCmFP4-STRIX-embF16-imatrix-headQ6.gguf \ | |
| --alias obliterated-27b-mtp \ | |
| --host 0.0.0.0 \ | |
| --port 8080 \ | |
| -dev Vulkan0 \ | |
| -ngl 999 \ | |
| -fa on \ | |
| -c 262144 \ | |
| -b 2048 \ | |
| -ub 256 \ | |
| -t 16 \ | |
| -tb 16 \ | |
| -ctk f16 \ | |
| -ctv f16 \ | |
| -cpent 256 \ | |
| -ctxcp 32 \ | |
| --cache-reuse 256 \ | |
| --cache-ram 65536 \ | |
| --temp 0.6 \ | |
| --top-p 0.95 \ | |
| --top-k 20 \ | |
| --min-p 0.0 \ | |
| --spec-type draft-mtp \ | |
| --spec-draft-device Vulkan0 \ | |
| --spec-draft-ngl all \ | |
| --spec-draft-type-k f16 \ | |
| --spec-draft-type-v f16 \ | |
| --spec-draft-n-max 5 \ | |
| --spec-draft-n-min 0 \ | |
| --spec-draft-p-min 0.0 \ | |
| --spec-draft-p-split 0.10 \ | |
| --chat-template-file chat_template.jinja \ | |
| --reasoning on \ | |
| --reasoning-format deepseek \ | |
| --chat-template-kwargs '{"preserve_thinking": true}' \ | |
| --jinja \ | |
| --parallel 1 \ | |
| --metrics \ | |
| --no-mmap \ | |
| --mmproj mmproj-F32.gguf \ | |
| --image-min-tokens 1024 | |
| ``` | |
| The **last two lines enable vision** — the `mmproj-F32.gguf` Qwen3-VL projector is **bundled in this repo** (`projection_dim 5120`); omit them for text-only. **`--image-min-tokens 1024` is required whenever `--mmproj` is set** (see §03). | |
| <div style="overflow:hidden; border-radius:0;"> | |
| <table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;"> | |
| <thead><tr> | |
| <th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px; width:40%;">Flag</th> | |
| <th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Function</th> | |
| </tr></thead> | |
| <tbody> | |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>HSA_OVERRIDE_GFX_VERSION=11.5.1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">treat the APU as gfx1151 (Strix Halo)</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>GGML_HIP_ENABLE_UNIFIED_MEMORY=1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">allow use of the full 128 GB unified memory</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-dev Vulkan0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">run on Vulkan — fastest backend for ROCmFP4 on Strix Halo</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ngl 999 · -fa on</code></td><td style="border:1px solid currentColor; padding:6px 10px;">offload all layers · flash attention</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-c 262144</code></td><td style="border:1px solid currentColor; padding:6px 10px;">context length (256K)</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-b 2048 · -ub 256 · -t/-tb 16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">prefill batch / micro-batch · CPU threads</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ctk f16 · -ctv f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">f16 KV cache — how we run it; drop to <code>q8_0</code>/<code>q4_0</code> to use less memory</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-cpent · -ctxcp · --cache-reuse · --cache-ram 65536</code></td><td style="border:1px solid currentColor; padding:6px 10px;">cross-turn KV checkpointing + 64 GB resident reuse cache</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">Qwen3.6 recommended sampling</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--spec-type draft-mtp · --spec-draft-n-max 5</code></td><td style="border:1px solid currentColor; padding:6px 10px;">grafted MTP head, self-speculative; draft depth 5</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--spec-draft-device Vulkan0 · -ngl all · type-k/v f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">draft head on Vulkan, fully offloaded, f16 KV (matches the main model)</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--chat-template-file chat_template.jinja</code></td><td style="border:1px solid currentColor; padding:6px 10px;">bundled froggeric template (tool calls + think-toggle + vision)</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--reasoning on --reasoning-format deepseek + kwargs {preserve_thinking:true}</code></td><td style="border:1px solid currentColor; padding:6px 10px;">thinking enabled, deepseek-style parsing; keep cross-turn cache</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--jinja --parallel 1 --metrics --no-mmap</code></td><td style="border:1px solid currentColor; padding:6px 10px;">apply template · single slot · metrics · weights in RAM</td></tr> | |
| </tbody> | |
| </table> | |
| </div> | |
| <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">03</span> · VISION</div> | |
| Qwen3-VL lineage — vision works via the **bundled `mmproj-F32.gguf`** projector with `--mmproj` (same LLM GGUF, no separate vision model). It's the Qwen3-VL projector (`projection_dim 5120`, matches this model's hidden size), shipped in this repo: | |
| ```bash | |
| --mmproj mmproj-F32.gguf \ | |
| --image-min-tokens 1024 # REQUIRED — Qwen-VL needs >=1024 image tokens or it misreads fine detail | |
| ``` | |
| Without `--image-min-tokens 1024` the server feeds too few image tokens and the model **describes images incorrectly** (right gist, wrong detail — the server even logs a warning at load). Verified on this model: a code label misread at default tokens read correctly once the flag was set. | |
| <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.85;"> | |
| <b>NOTE //</b> thinking model → for one-shot image Q&A use <code><|think_off|></code> (the bundled <code>chat_template.jinja</code>) or allow enough tokens, else the answer can come back empty. With <code>--mmproj</code> loaded the server disables the <code>--cache-reuse</code> feature (it logs <i>"cache_reuse is not supported by multimodal"</i>); whether ordinary cross-turn caching still helps with vision isn't something we've benchmarked. | |
| </div> | |
| <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">04</span> · PERFORMANCE & QUALITY</div> | |
| **This is the best speed/quality balance in ROCmFP4 — by design, not the absolute fastest.** The two quality levers that are actually *felt* — genuine **f16 token embeddings** and a **Q6_K output head** — sit on the fast single-scale `q4_0_rocmfp4_fast` body. The alternatives we'd otherwise reach for (an all-dual-scale body, selective higher-precision tensors) buy a KL improvement that sits *inside the measurement noise* while costing decode speed — so the fast body is the right point. A leaner Q5-embedding build is a few tok/s faster but degrades the one lever you notice; we keep full f16 embeddings. | |
| We ran this exact sweep in full on the sibling **[Qwen3.6-27B base card](https://huggingface.co/plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF)** — every rocmfp4 lever measured by **KL divergence vs the BF16 reference** plus `llama-bench` decode. The frontier there is the same shape it is here: the fast body + f16 emb + Q6 head is the balance point, all-dual-scale and selective higher-precision land inside the noise, and the dynamic K-quant is the fidelity ceiling that rocmfp4's FP4 can't out-allocate. See that card for the numbers and the full experiments table; OBLITERATED follows the same recipe. | |
| <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.9;"> | |
| <b>WANT MAXIMUM FIDELITY INSTEAD OF SPEED?</b> There's no Unsloth UD-quant of this abliterated model, so for the last bit of fidelity grab a <b>Q6_K / Q8 GGUF of the base abliterated model</b> from <a href="https://huggingface.co/OBLITERATUS/Qwen3.6-27B-OBLITERATED"><b>OBLITERATUS/Qwen3.6-27B-OBLITERATED</b></a>. Those higher-bit GGUFs run on this same fork — trading lower KL for slower decode. We optimize for throughput in ROCmFP4; if you want fidelity over speed, that's the one to grab. | |
| </div> | |
| **Grafted MTP head — measured.** OBLITERATED ships **no** MTP head, so we transplanted a `nextn` block from a **Qwen3.6-27B-MTP BF16 donor** (output-lossless — the draft head only affects *speed*, never the tokens). Because OBLITERATED is abliterated *from* that same base, the borrowed head tracks it closely. Measured on the ROCmFP4 fork (`llama-server --spec-type draft-mtp`, f16 KV on target **and** drafter, 4 prompt types, temp 0.6): | |
| <div style="overflow:hidden; border-radius:0;"> | |
| <table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;"> | |
| <thead><tr> | |
| <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Content</th> | |
| <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Decode t/s (MTP)</th> | |
| <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">t/s (no MTP)</th> | |
| <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Draft acceptance</th> | |
| </tr></thead> | |
| <tbody> | |
| <tr><td style="border:1px solid currentColor; padding:7px 10px;">math / reasoning</td><td style="border:1px solid currentColor; padding:7px 10px;">35.1</td><td style="border:1px solid currentColor; padding:7px 10px;">14.0</td><td style="border:1px solid currentColor; padding:7px 10px;">88.1%</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:7px 10px;">technical</td><td style="border:1px solid currentColor; padding:7px 10px;">29.5</td><td style="border:1px solid currentColor; padding:7px 10px;">14.1</td><td style="border:1px solid currentColor; padding:7px 10px;">69.1%</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:7px 10px;">code</td><td style="border:1px solid currentColor; padding:7px 10px;">28.8</td><td style="border:1px solid currentColor; padding:7px 10px;">14.1</td><td style="border:1px solid currentColor; padding:7px 10px;">67.5%</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:7px 10px;">prose / creative</td><td style="border:1px solid currentColor; padding:7px 10px;">24.6</td><td style="border:1px solid currentColor; padding:7px 10px;">14.0</td><td style="border:1px solid currentColor; padding:7px 10px;">52.2%</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">average</td><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">29.5</td><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">14.1</td><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">67.7%</td></tr> | |
| </tbody> | |
| </table> | |
| </div> | |
| **~2.1× faster decode** (14.1 → 29.5 t/s) at **67.7%** token acceptance — high on structured/predictable text, lower on freeform prose, the profile of a borrowed (not natively-trained) head. KV is f16 on both the main model and the draft; the draft head is kept at 4-bit ROCmFP4. | |
| **The imatrix WINS here — measured.** Quantized **with** an importance matrix from a public general+code calibration mix (Kalomaze `groups_merged` + froggeric `code`/`technical`, via [`froggeric/imatrix`](https://huggingface.co/datasets/froggeric/imatrix)). Measured by KL-divergence + perplexity vs the **true BF16** on a held-out general slice, imatrix vs no-imatrix: | |
| <div style="overflow:hidden; border-radius:0;"> | |
| <table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;"> | |
| <thead><tr> | |
| <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Metric (vs BF16, held-out general)</th> | |
| <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">No-imatrix</th> | |
| <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Imatrix</th> | |
| <th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Change</th> | |
| </tr></thead> | |
| <tbody> | |
| <tr><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">Perplexity</td><td style="border:1px solid currentColor; padding:7px 10px;">+3.08%</td><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">+2.58%</td><td style="border:1px solid currentColor; padding:7px 10px;">recovers ~16% of the 4-bit loss</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">Median KLD</td><td style="border:1px solid currentColor; padding:7px 10px;">0.02070</td><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">0.01843</td><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">−11%</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:7px 10px;">Mean KLD</td><td style="border:1px solid currentColor; padding:7px 10px;">0.04239</td><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">0.03961</td><td style="border:1px solid currentColor; padding:7px 10px;">−6.6%</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:7px 10px;">99th-pct KLD</td><td style="border:1px solid currentColor; padding:7px 10px;">0.3729</td><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">0.3298</td><td style="border:1px solid currentColor; padding:7px 10px;">−12%</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:7px 10px;">RMS Δp</td><td style="border:1px solid currentColor; padding:7px 10px;">6.37%</td><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">6.13%</td><td style="border:1px solid currentColor; padding:7px 10px;">−3.7%</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">Same top token as BF16</td><td style="border:1px solid currentColor; padding:7px 10px;">90.59%</td><td style="border:1px solid currentColor; padding:7px 10px; font-weight:700;">91.06%</td><td style="border:1px solid currentColor; padding:7px 10px;">+0.47 pp</td></tr> | |
| </tbody> | |
| </table> | |
| </div> | |
| For OBLITERATED the imatrix is a **clean win** on every robust metric — it behaves like its base Qwen3.6-27B, not like the dense [Qwopus-Coder](https://huggingface.co/plunderstruck/Qwopus3.6-27B-Coder-MTP-ROCmFP4-GGUF) (where the same recipe *worsened* code-PPL — but that was a *code* metric; this is a general model measured on general text). Always measure; we did. | |
| <div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:14px 0; opacity:0.85;"> | |
| <b>NOTE //</b> Quality scope: the KL/PPL above is a fidelity-vs-BF16 measurement on ~20 K tokens of held-out general text, <b>not</b> an absolute benchmark. The MTP head is <i>borrowed</i> (not trained on this model) — output-lossless, affecting speed only. Experimental research build for AMD Strix Halo — hardware/driver/prompt-sensitive. <b>Not</b> native FP4 tensor-core execution. | |
| </div> | |
| <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">05</span> · BUILD (REPRODUCIBLE)</div> | |
| ```bash | |
| # 0) convert the abliterated safetensors -> BF16 GGUF | |
| python convert_hf_to_gguf.py OBLITERATED-27b/ --outtype bf16 --outfile OBLITERATED-BF16.gguf | |
| # 1) ** GOTCHA ** OBLITERATED's config declares `mtp_num_hidden_layers: 1` but ships NO MTP weights, | |
| # so the convert labels it block_count=65 with only 64 real layers -> "missing tensor blk.64.*" on load. | |
| # The real model is 64 layers; the donor's nextn becomes the new blk.64. | |
| # 2) graft a Qwen3.6-27B nextn head onto blk.64, then set block_count = 65 | |
| python inject_mtp_40b.py --target OBLITERATED-BF16.gguf --donor Qwen3.6-27B-MTP-BF16.gguf \ | |
| --output OBLITERATED-MTP-BF16.gguf --source-layer 64 --dest-layer 64 | |
| python gguf_set_metadata.py OBLITERATED-MTP-BF16.gguf qwen35.block_count 65 --force | |
| # 3) imatrix on the grafted BF16, then quant -> ROCmFP4 with genuine f16 embeddings | |
| llama-imatrix -m OBLITERATED-MTP-BF16.gguf -f general+code-calib.txt -o obliterated.imatrix -c 512 -ngl 999 | |
| llama-quantize --token-embedding-type f16 --imatrix obliterated.imatrix \ | |
| OBLITERATED-MTP-BF16.gguf Qwen3.6-27B-OBLITERATED-MTP-ROCmFP4-STRIX-embF16-imatrix.gguf Q4_0_ROCMFP4_STRIX | |
| # headQ6 variant adds the Q6_K output head | |
| llama-quantize --token-embedding-type f16 --output-tensor-type q6_K --imatrix obliterated.imatrix \ | |
| OBLITERATED-MTP-BF16.gguf Qwen3.6-27B-OBLITERATED-MTP-ROCmFP4-STRIX-embF16-imatrix-headQ6.gguf Q4_0_ROCMFP4_STRIX | |
| ``` | |
| > Experimental research build for AMD Strix Halo — hardware/driver/prompt-sensitive, may not reproduce elsewhere. Not native FP4 tensor-core execution. | |
| <div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">06</span> · LINEAGE & CREDITS</div> | |
| ``` | |
| this ROCmFP4 quant ──quantized──▶ OBLITERATUS/Qwen3.6-27B-OBLITERATED ──abliterated──▶ Qwen/Qwen3.6-27B | |
| ``` | |
| A 4-bit Strix-Halo quant of OBLITERATUS's abliterated 27B. The abliteration (refusal-direction removal + source interpolation) is upstream work; we add the ROCmFP4 quant, genuine f16 embeddings, and a grafted MTP head. | |
| <div style="overflow:hidden; border-radius:0;"> | |
| <table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;"> | |
| <tbody> | |
| <tr><td style="border:1px solid currentColor; padding:8px 11px; width:26%;">BASE MODEL</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/OBLITERATUS/Qwen3.6-27B-OBLITERATED">OBLITERATUS/Qwen3.6-27B-OBLITERATED</a> (Apache-2.0) · abliterated from <a href="https://huggingface.co/Qwen/Qwen3.6-27B">Qwen/Qwen3.6-27B</a> (Qwen team)</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:8px 11px;">MTP DONOR</td><td style="border:1px solid currentColor; padding:8px 11px;">a Qwen3.6-27B-MTP <code>nextn</code> head (graft is output-lossless)</td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:8px 11px;">CALIBRATION</td><td style="border:1px solid currentColor; padding:8px 11px;">Kalomaze <code>groups_merged</code> + froggeric <code>code</code>/<code>technical</code> via <a href="https://huggingface.co/datasets/froggeric/imatrix">froggeric/imatrix</a></td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:8px 11px;">CHAT TEMPLATE</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates">froggeric/Qwen-Fixed-Chat-Templates</a></td></tr> | |
| <tr><td style="border:1px solid currentColor; padding:8px 11px;">FORMAT + RUNTIME</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://github.com/charlie12345/ROCmFPX">charlie12345/ROCmFPX</a> (llama.cpp, MIT)</td></tr> | |
| </tbody> | |
| </table> | |
| </div> | |
| **Sibling ROCmFP4 Strix Halo models** — [Qwen3.6-27B-MTP](https://huggingface.co/plunderstruck/Qwen3.6-27B-MTP-ROCmFP4-GGUF) · [Qwen3.6-35B-A3B-MTP](https://huggingface.co/plunderstruck/Qwen3.6-35B-A3B-MTP-ROCmFP4-GGUF) · [Qwopus3.6-27B-Coder-MTP](https://huggingface.co/plunderstruck/Qwopus3.6-27B-Coder-MTP-ROCmFP4-GGUF) · [Qwen3.6-40B-Deckard-MTP](https://huggingface.co/plunderstruck/Qwen3.6-40B-Deckard-MTP-ROCmFP4-GGUF) · [Qwen3-Coder-Next](https://huggingface.co/plunderstruck/Qwen3-Coder-Next-ROCmFP4-GGUF) · [Nex-N2-mini](https://huggingface.co/plunderstruck/Nex-N2-mini-ROCmFP4-GGUF) | |
| *Derivative quantization — verify the base model's license before redistribution / use.* | |