plunderstruck's picture
WIP card: ROCmFP4 quant coming soon (build in progress)
4867049 verified
|
Raw History Blame
4.28 kB
metadata
base_model: OBLITERATUS/Qwen3.6-27B-OBLITERATED
base_model_relation: quantized
license: apache-2.0
library_name: gguf
tags:
  - gguf
  - rocmfp4
  - qwen3.6
  - obliterated
  - abliterated
  - uncensored
  - 27b
  - mtp
  - speculative-decoding
  - strix-halo
  - amd
  - rocm
  - vulkan
language:
  - en

Qwen3.6-27B-OBLITERATED-MTP — ROCmFP4 STRIX

🚧 Work in progress — files not uploaded yet

This card is up so you know it's coming. The GGUFs aren't here yet — the BF16 source is downloading and the quant/measurement pipeline is queued behind some other uploads. Check back; this notice gets replaced with the files + measured numbers when it's done.

Experimental AMD Strix Halo (gfx1151) quant of OBLITERATUS/Qwen3.6-27B-OBLITERATED — an abliterated / uncensored Qwen3.6-27B (refusal direction removed, with source-weight interpolation to retain capability) — in the custom ROCmFP4 4-bit format, with an MTP / next-token-prediction head grafted in for self-speculative decoding on a single APU.

⚠️ Ignore HuggingFace's auto-detected quant badge ("F16"/16-bit) — it's wrong. HF can't read the custom ROCmFP4 tensor types and mislabels the file by its f16 embeddings. These will be ~4.5 bpw 4-bit ROCmFP4 files, not 16-bit. Pick by filename in Files and versions.

Requires the ROCmFP4 fork (public) — not stock llama.cpp

Uses the ROCmFP4 tensor types (q4_0_rocmfp4, q4_0_rocmfp4_fast). Stock llama.cpp, LM Studio, Ollama, etc. cannot load it. Build/run with charlie12345/rocmfp4-llama (mtp-rocmfp4-strix).

Lineage

this ROCmFP4 quant  ──quantized──▶  OBLITERATUS/Qwen3.6-27B-OBLITERATED  ──abliterated──▶  Qwen/Qwen3.6-27B

It's a 4-bit Strix-Halo quant of OBLITERATUS's abliterated 27B. The abliteration (refusal-direction removal + source interpolation) is upstream work; we only do the ROCmFP4 quant + MTP graft.

Planned recipe (what these will be)

  • Genuine f16 token embeddings — quantized from the BF16 safetensors, so the f16 embeddings are real f16 (not Q8→f16 "fake-f16").
  • Grafted MTP head — a nextn head transplanted from a Qwen3.6-27B-MTP BF16 donor (output-lossless; it only affects draft speed), so it runs self-speculative on the fork.
  • Two output-head variants — a base (4-bit head, fastest) and a Q6_K head (a notch more faithful).
  • imatrix: measured, then decided. This is a dense Qwen3.6-27B — the same architecture class as our Qwopus-Coder, the one model where a code-weighted imatrix worsened perplexity. So we'll build both and pick by measured KL + perplexity vs the BF16, rather than assume. Whatever ships, the card will say which won and show the numbers.

Status

🚧 Building. No files yet, and no quality measurement yet — nothing here is a quality claim. When it lands it'll carry the same honest scope as our other cards: a KL/PPL fidelity comparison vs BF16, not an absolute benchmark.

Sibling ROCmFP4 Strix Halo models

Credits & license