Qwen3.8-Flash-Next — Strix Halo Halogen GGUF

This is a custom checkpoint for halogen-flash-server, a Qwen4-Exp runtime for AMD Strix Halo (gfx1151). The .hgn is halogen's own resident-arena format and does not load in llama.cpp, LM Studio, or Ollama; the GGUF is tested with halogen 0.12.1 (see Runtime). Loading with an incompatible runtime may fail or produce invalid output.

An abliterated, importance-matrix-calibrated quantization of the official Qwen/Qwen3.8-Flash-Next release, with its MTP draft head. Derived through peonist-ai/halogen-qwen3.8-flash-next, halogen's packaging of the Qwen base. Two servable artifacts ship here: a GGUF (IQ4_XS experts) and the native .hgn resident-arena checkpoint the engine prefers.

Artifacts

Verify any local copy against the SHA-256 values below.

file bytes GiB SHA-256
Qwen3.8-Flash-Next-abliterated-IQ4_XS.gguf 97,615,265,184 90.9113 9bfc815a5f5e3045de7d8513ab0dbb66d174108a2b27d7e7a8d58ae53af9b15f
Qwen3.8-Flash-Next-abliterated.hgn 124,429,843,520 115.8843 bb542c77e9a78062637af4c82dd2e20debae04d85db9bd562db12dd64b4ebba7
qwen38-flash-next-mtp.hgn 1,523,566,720 1.4189 0f50e9626df98168e7c6e0cc264e2a92b5184dd885a175a06628d979b5edceeb
tokenizer/ — — —

The .hgn is 115.88 GiB on disk, including a 47.68 GiB FP8 n-gram lookup table that is demand-paged rather than pinned in full. File size is not resident memory. At the tested configuration (262,144-token context, 524,288 pooled KV positions, four slots, vision enabled), consult the engine's startup memory breakdown. GPU-pinned file pages can make Linux MemAvailable overstate usable headroom. GTT is shared host memory, not additional RAM.

The GGUF carries an IQ4_NL n-gram table (26.82 GiB) for compatibility with halogen's GGUF importer. The native .hgn carries FP8, converted directly from the BF16 source. The two table representations are not equivalent. Use the .hgn for the FP8 table. A vision tower is not included; image input requires the tower from the halogen release.

Quantization recipe

Importance-matrix-calibrated llama-quantize with per-tensor overrides for the trunk and experts, plus direct BF16-to-FP8 conversion of the native table:

tensors type
routed expert gate/up matrices IQ4_XS (imatrix)
routed expert down matrices IQ4_NL (48 shape fallbacks from IQ4_XS)
shared expert, dense, attention projections, hyper-connection mixers, output IQ4_NL
token embedding Q8_0
PLE / n-gram table in GGUF IQ4_NL
PLE / n-gram table in native HGN FP8 E4M3FN, one global FP32 scale
preserved convolution, indexer, norms and other small tensors BF16 or F32, per overrides

GGUF counts: 546 IQ4_NL, 96 IQ4_XS, 193 BF16, 388 F32, 1 Q8_0; zero F16 fallbacks. See quantization-overrides.txt and quantization.json.

The native table is quantized directly from all 128 BF16 embedding shards: scale = max(abs(weight)) / 448, round-to-nearest with ties-to-even, E4M3FN bytes followed by the FP32 scale. All 51,200,245,760 values are retained. This conversion does not expand a 4-bit table back to 8 bits.

The importance matrix is collected on the base model and applied to the edited weights — an approximation, since the edit shifts the activations the matrix described.

The abliteration band

The refusal-direction edit uses a rank-1 residual-writer projection (W' = W − λ · r rᵀ W) on every write matrix that adds to the residual stream: all routed experts' down_proj, attention output (o_proj / out_proj / attn_output), and the shared-expert down. λ_attn = 3.5, applied across all 48 layers.

The refusal direction r is captured per layer from block-output activations over 48 harmful and 48 harmless prompts (seed 42). Shipped with HALOGEN_CK_OVERLAY=none as the floor, because the stock exclusion overlay would revert 96 of the edited tensors.

Behavioural validation

Run on the served FP8-table .hgn:

set requirement result
arithmetic exact integer answer PASS
instruction following requested three-word answer PASS
Python code requested function and return statement PASS
image input correct fixture color PASS

These are bounded integration checks. Refusal-rate and over-trigger evaluations have not been run on this artifact; no behavioural removal rate is certified here.

Structural validation

The .hgn carries 1198 tensors, including its MTP head. Validation checks every tensor payload checksum, bounds, alignment and non-overlap. SHA-256 comparisons verify all 1,197 non-table tensor payloads are identical to the validated trunk used to assemble this checkpoint.

The table uses native HGN type fp8g, shape [128, 2500012, 160], and occupies 51,200,245,764 bytes, including its scale. Separate readback checked sampled rows against BF16 across every shard. The complete table is byte-identical to the stock FP8 table after conversion from this release's BF16 source.

Halogen 0.12.1 loaded the native checkpoint and passed serving checks. Its GGUF converter requires IQ4_NL for this tensor; converting the accompanying GGUF produces an IQ4-table checkpoint, not this FP8-table HGN.

The MTP draft head

The checkpoint carries its own MTP head — a full Flash-Next layer — used for speculative decoding when serving the .hgn. qwen38-flash-next-mtp.hgn is the same head packaged standalone, for the bring-your-own-GGUF path where the GGUF has no head of its own.

Quantization quality — not yet characterized

A seeded sample of 368,640 embedding values across all 128 shards measured relative reconstruction RMSE of 2.6551% for FP8 and 7.6200% for IQ4_NL against BF16. This measures table reconstruction error, not end-to-end model quality.

No perplexity, KL-divergence, broad capability benchmark or controlled end-to-end quality comparison has been run on this artifact.

Runtime

Verified with halogen-flash-server 0.12.1. Other versions are unverified. The engine is gfx1151-only. Its native loader accepts this FP8 table; its GGUF importer requires the accompanying IQ4_NL table.

Serve the .hgn (recommended, verified):

hf download otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF --local-dir ~/halogen-models

podman run --rm -p 8731:8731 \
  --device /dev/kfd --device /dev/dri --group-add keep-groups \
  --ipc=host --ulimit memlock=-1:-1 \
  -e HALOGEN_CHECKPOINT=/models/Qwen3.8-Flash-Next-abliterated.hgn \
  -e HALOGEN_CK_OVERLAY=none \
  -v ~/halogen-models:/models:ro \
  ghcr.io/peonist-ai/halogen-flash-server:0.12.1

On Docker (not Podman), replace --group-add keep-groups with the numeric render/video GIDs, e.g. --group-add 39 --group-add 105. An OpenAI-compatible endpoint comes up on :8731 (/v1/chat/completions, /v1/completions, /v1/models, /v1/responses):

curl http://localhost:8731/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"halogen-qwen3.8-flash-next",
       "messages":[{"role":"user","content":"Hello"}]}'

Do not place an overlay sidecar (<checkpoint>.overlay.hgn) beside this checkpoint — this repo ships none on purpose, and the stock overlay reverts the edit. To serve the GGUF instead, 0.12.1 opens it and repacks into RAM; that path needs the GGUF, tokenizer/, and qwen38-flash-next-mtp.hgn. Download the native .hgn above to serve the FP8 table.

Give the server a machine of its own: at the defaults it holds most of a 125 GiB Strix Halo host, and under memory pressure it can stall for minutes at full CPU with no output — host-memory compaction, not a crash.

Provenance

Qwen/Qwen3.8-Flash-Next
  (via peonist-ai/halogen-qwen3.8-flash-next, halogen's packaging)
  -> rank-1 refusal projection on residual writers, lambda_attn 3.5
     experts' down_proj + attention output + shared-expert down, 48 layers
     directions from 48 harmful / 48 harmless block-outputs (seed 42)
  -> edited BF16 source
  -> imatrix-calibrated IQ4_XS/IQ4_NL experts and IQ4_NL trunk
  -> GGUF with IQ4_NL n-gram table for importer compatibility
  -> native HGN trunk and MTP assembly
     + direct BF16 -> FP8 E4M3FN n-gram table conversion
  -> full payload validation and served smoke checks (overlay=none)

License

This artifact inherits the Qwen model license (Qwen Community License 1.0). See the base model card and LICENSE for its terms.

Downloads last month
400
GGUF
Model size
177B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF

Quantized
(304)
this model