Instructions to use otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS # Run inference directly in the terminal: llama cli -hf otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS # Run inference directly in the terminal: llama cli -hf otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS # Run inference directly in the terminal: ./llama-cli -hf otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS # Run inference directly in the terminal: ./build/bin/llama-cli -hf otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS
Use Docker
docker model run hf.co/otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS
- LM Studio
- Jan
- vLLM
How to use otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS
- Ollama
How to use otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF with Ollama:
ollama run hf.co/otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS
- Unsloth Desktop
- Pi
How to use otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF with Docker Model Runner:
docker model run hf.co/otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS
- Lemonade
How to use otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Next-Strix-Halo-GGUF-IQ4_XS
List all available models
lemonade list
- Hermes Agent
How to use otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF:IQ4_XS" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8-Flash-Next — Strix Halo Halogen GGUF
This is a custom checkpoint for halogen-flash-server, a Qwen4-Exp runtime for AMD Strix Halo (
gfx1151). The.hgnis halogen's own resident-arena format and does not load inllama.cpp, LM Studio, or Ollama; the GGUF is tested with halogen 0.12.1 (see Runtime). Loading with an incompatible runtime may fail or produce invalid output.
An abliterated, importance-matrix-calibrated quantization of the official
Qwen/Qwen3.8-Flash-Next release, with its MTP draft head. Derived through
peonist-ai/halogen-qwen3.8-flash-next,
halogen's packaging of the Qwen base. Two servable artifacts ship here: a GGUF
(IQ4_XS experts) and the native .hgn resident-arena checkpoint the engine prefers.
Artifacts
Verify any local copy against the SHA-256 values below.
| file | bytes | GiB | SHA-256 |
|---|---|---|---|
Qwen3.8-Flash-Next-abliterated-IQ4_XS.gguf |
97,615,265,184 |
90.9113 |
9bfc815a5f5e3045de7d8513ab0dbb66d174108a2b27d7e7a8d58ae53af9b15f |
Qwen3.8-Flash-Next-abliterated.hgn |
124,429,843,520 |
115.8843 |
bb542c77e9a78062637af4c82dd2e20debae04d85db9bd562db12dd64b4ebba7 |
qwen38-flash-next-mtp.hgn |
1,523,566,720 |
1.4189 |
0f50e9626df98168e7c6e0cc264e2a92b5184dd885a175a06628d979b5edceeb |
tokenizer/ |
— | — | — |
The .hgn is 115.88 GiB on disk, including a 47.68 GiB FP8
n-gram lookup table that is demand-paged rather than pinned in full. File size
is not resident memory. At the tested configuration (262,144-token context,
524,288 pooled KV positions, four slots, vision enabled), consult the engine's
startup memory breakdown. GPU-pinned file pages can make Linux MemAvailable
overstate usable headroom. GTT is shared host memory, not additional RAM.
The GGUF carries an IQ4_NL n-gram table (26.82 GiB) for compatibility with
halogen's GGUF importer. The native .hgn carries FP8, converted directly
from the BF16 source. The two table representations are not equivalent.
Use the .hgn for the FP8 table. A vision tower is not included;
image input requires the tower from the halogen release.
Quantization recipe
Importance-matrix-calibrated llama-quantize with per-tensor overrides for
the trunk and experts, plus direct BF16-to-FP8 conversion of the native table:
| tensors | type |
|---|---|
| routed expert gate/up matrices | IQ4_XS (imatrix) |
| routed expert down matrices | IQ4_NL (48 shape fallbacks from IQ4_XS) |
| shared expert, dense, attention projections, hyper-connection mixers, output | IQ4_NL |
| token embedding | Q8_0 |
| PLE / n-gram table in GGUF | IQ4_NL |
| PLE / n-gram table in native HGN | FP8 E4M3FN, one global FP32 scale |
| preserved convolution, indexer, norms and other small tensors | BF16 or F32, per overrides |
GGUF counts: 546 IQ4_NL, 96 IQ4_XS, 193 BF16, 388 F32, 1 Q8_0;
zero F16 fallbacks. See quantization-overrides.txt and quantization.json.
The native table is quantized directly from all 128 BF16 embedding shards:
scale = max(abs(weight)) / 448, round-to-nearest with ties-to-even, E4M3FN
bytes followed by the FP32 scale. All 51,200,245,760 values are retained.
This conversion does not expand a 4-bit table back to 8 bits.
The importance matrix is collected on the base model and applied to the edited weights — an approximation, since the edit shifts the activations the matrix described.
The abliteration band
The refusal-direction edit uses a rank-1 residual-writer projection (W' = W − λ · r rᵀ W)
on every write matrix that adds to the residual stream: all routed experts'
down_proj, attention output (o_proj / out_proj / attn_output), and the
shared-expert down. λ_attn = 3.5, applied across all 48 layers.
The refusal direction r is captured per layer from block-output activations
over 48 harmful and 48 harmless prompts (seed 42). Shipped with
HALOGEN_CK_OVERLAY=none as the floor, because the stock exclusion overlay would
revert 96 of the edited tensors.
Behavioural validation
Run on the served FP8-table .hgn:
| set | requirement | result |
|---|---|---|
| arithmetic | exact integer answer | PASS |
| instruction following | requested three-word answer | PASS |
| Python code | requested function and return statement | PASS |
| image input | correct fixture color | PASS |
These are bounded integration checks. Refusal-rate and over-trigger evaluations have not been run on this artifact; no behavioural removal rate is certified here.
Structural validation
The .hgn carries 1198 tensors, including its MTP head. Validation checks
every tensor payload checksum, bounds, alignment and non-overlap. SHA-256
comparisons verify all 1,197 non-table tensor payloads are identical to the
validated trunk used to assemble this checkpoint.
The table uses native HGN type fp8g, shape [128, 2500012, 160], and occupies
51,200,245,764 bytes, including its scale. Separate readback checked sampled
rows against BF16 across every shard. The complete table is byte-identical to
the stock FP8 table after conversion from this release's BF16 source.
Halogen 0.12.1 loaded the native checkpoint and passed serving checks. Its GGUF converter requires IQ4_NL for this tensor; converting the accompanying GGUF produces an IQ4-table checkpoint, not this FP8-table HGN.
The MTP draft head
The checkpoint carries its own MTP head — a full Flash-Next layer — used for
speculative decoding when serving the .hgn. qwen38-flash-next-mtp.hgn is the
same head packaged standalone, for the bring-your-own-GGUF path where the GGUF
has no head of its own.
Quantization quality — not yet characterized
A seeded sample of 368,640 embedding values across all 128 shards measured relative reconstruction RMSE of 2.6551% for FP8 and 7.6200% for IQ4_NL against BF16. This measures table reconstruction error, not end-to-end model quality.
No perplexity, KL-divergence, broad capability benchmark or controlled end-to-end quality comparison has been run on this artifact.
Runtime
Verified with halogen-flash-server 0.12.1. Other versions are unverified.
The engine is gfx1151-only. Its native loader accepts this FP8 table; its
GGUF importer requires the accompanying IQ4_NL table.
Serve the .hgn (recommended, verified):
hf download otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF --local-dir ~/halogen-models
podman run --rm -p 8731:8731 \
--device /dev/kfd --device /dev/dri --group-add keep-groups \
--ipc=host --ulimit memlock=-1:-1 \
-e HALOGEN_CHECKPOINT=/models/Qwen3.8-Flash-Next-abliterated.hgn \
-e HALOGEN_CK_OVERLAY=none \
-v ~/halogen-models:/models:ro \
ghcr.io/peonist-ai/halogen-flash-server:0.12.1
On Docker (not Podman), replace --group-add keep-groups with the numeric
render/video GIDs, e.g. --group-add 39 --group-add 105. An OpenAI-compatible
endpoint comes up on :8731 (/v1/chat/completions, /v1/completions,
/v1/models, /v1/responses):
curl http://localhost:8731/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"halogen-qwen3.8-flash-next",
"messages":[{"role":"user","content":"Hello"}]}'
Do not place an overlay sidecar (<checkpoint>.overlay.hgn) beside this
checkpoint — this repo ships none on purpose, and the stock overlay reverts the
edit. To serve the GGUF instead, 0.12.1 opens it and repacks into RAM; that path
needs the GGUF, tokenizer/, and qwen38-flash-next-mtp.hgn. Download the native
.hgn above to serve the FP8 table.
Give the server a machine of its own: at the defaults it holds most of a 125 GiB Strix Halo host, and under memory pressure it can stall for minutes at full CPU with no output — host-memory compaction, not a crash.
Provenance
Qwen/Qwen3.8-Flash-Next
(via peonist-ai/halogen-qwen3.8-flash-next, halogen's packaging)
-> rank-1 refusal projection on residual writers, lambda_attn 3.5
experts' down_proj + attention output + shared-expert down, 48 layers
directions from 48 harmful / 48 harmless block-outputs (seed 42)
-> edited BF16 source
-> imatrix-calibrated IQ4_XS/IQ4_NL experts and IQ4_NL trunk
-> GGUF with IQ4_NL n-gram table for importer compatibility
-> native HGN trunk and MTP assembly
+ direct BF16 -> FP8 E4M3FN n-gram table conversion
-> full payload validation and served smoke checks (overlay=none)
License
This artifact inherits the Qwen model license (Qwen Community License 1.0). See
the base model card and LICENSE for its terms.
- Downloads last month
- 400
4-bit
Model tree for otheru/Qwen3.8-Flash-Next-Strix-Halo-GGUF
Base model
Qwen/Qwen3.8-Flash-Next