How to use from
OpenClaw
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf him0413/Qwen3.8-Flash-Next-Uncensored-Q4_K_M-MTP:Q4_K_M
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "him0413/Qwen3.8-Flash-Next-Uncensored-Q4_K_M-MTP:Q4_K_M" \
  --custom-provider-id llama-cpp \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

Qwen3.8-Flash-Next-Uncensored Q4_K_M (Integrated MTP)

An integrated MTP (Multi-Token Prediction) GGUF build of the abliterated Qwen3.8-Flash-Next-Uncensored model, quantized to Q4_K_M with the MTP draft head merged into the same 4-shard split — no sidecar file needed.

Model Details

Property Value
Base model Qwen/Qwen3.8-Flash-Next
Abliteration orcarouter/Qwen3.8-Flash-Next-Uncensored
Quant Q4_K_M (main trunk) + Q4_K_M (MTP head)
Split 4 shards (3 trunk + 1 MTP), integrated
Total size ~113.7 GiB
Context 262K native
Architecture qwen4exp (Gated DeltaNet + QSA + HyperConnections + PLE)
MTP 1 draft layer, integrated as blk.48

Usage (llama.cpp with qwen4exp support)

./build-vulkan/bin/llama-server \
  --model Qwen3.8-Flash-Next-Uncensored-Q4_K_M-MTP-00001-of-00004.gguf \
  --flash-attn on \
  --spec-type draft-mtp \
  --spec-draft-adaptive \
  --spec-draft-n-min 0 \
  --spec-draft-n-max 7 \
  --spec-draft-p-min 0.75

What's inside

This repo contains the MTP head tensors fused into the main GGUF:

  • blk.48.nextn.* — MTP embedding/hidden projections + norms
  • blk.48.nextn.hc_head_* — draft head output mixer
  • Full 48-layer trunk with nextn_predict_layers = 1

⚠️ Disclaimer

This model is an abliterated (refusal-removed) build. It will comply with harmful, unethical, or illegal requests the original Qwen3.8-Flash-Next would refuse. Released strictly for legitimate research — interpretability, AI-safety / refusal-mechanism study, red-teaming, and robustness evaluation. You assume full responsibility for how you use it and everything it generates; add your own safety and moderation layers before any deployment.

License

Qwen Community License 1.0 (see LICENSE), inherited from the base model Qwen/Qwen3.8-Flash-Next. Abliteration and quantization do not change the underlying license obligations.

Note: per the Qwen Community License, if you operate a Model as a Service or AI Work Assistant business commercially, you must obtain a separate license from Qwen before using this model or its derivatives for commercial purposes. See LICENSE for full terms.

Downloads last month
226
GGUF
Model size
180B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for him0413/Qwen3.8-Flash-Next-Uncensored-Q4_K_M-MTP

Quantized
(172)
this model