him0413's picture
Upload README.md with huggingface_hub
a8bccef verified
|
Raw
History Blame Contribute Delete
2.85 kB
metadata
license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model:
  - Qwen/Qwen3.8-Flash-Next
  - orcarouter/Qwen3.8-Flash-Next-Uncensored
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
language:
  - en
  - zh
tags:
  - qwen
  - qwen4
  - qwen3.8
  - flash-next
  - uncensored
  - abliterated
  - moe
  - gguf
  - llama.cpp
  - speculative
  - mtp
  - vision-language

Qwen3.8-Flash-Next-Uncensored Q4_K_M (Integrated MTP)

An integrated MTP (Multi-Token Prediction) GGUF build of the abliterated Qwen3.8-Flash-Next-Uncensored model, quantized to Q4_K_M with the MTP draft head merged into the same 4-shard split — no sidecar file needed.

Model Details

Property Value
Base model Qwen/Qwen3.8-Flash-Next
Abliteration orcarouter/Qwen3.8-Flash-Next-Uncensored
Quant Q4_K_M (main trunk) + Q4_K_M (MTP head)
Split 4 shards (3 trunk + 1 MTP), integrated
Total size ~113.7 GiB
Context 262K native
Architecture qwen4exp (Gated DeltaNet + QSA + HyperConnections + PLE)
MTP 1 draft layer, integrated as blk.48

Usage (llama.cpp with qwen4exp support)

./build-vulkan/bin/llama-server \
  --model Qwen3.8-Flash-Next-Uncensored-Q4_K_M-MTP-00001-of-00004.gguf \
  --flash-attn on \
  --spec-type draft-mtp \
  --spec-draft-adaptive \
  --spec-draft-n-min 0 \
  --spec-draft-n-max 7 \
  --spec-draft-p-min 0.75

What's inside

This repo contains the MTP head tensors fused into the main GGUF:

  • blk.48.nextn.* — MTP embedding/hidden projections + norms
  • blk.48.nextn.hc_head_* — draft head output mixer
  • Full 48-layer trunk with nextn_predict_layers = 1

⚠️ Disclaimer

This model is an abliterated (refusal-removed) build. It will comply with harmful, unethical, or illegal requests the original Qwen3.8-Flash-Next would refuse. Released strictly for legitimate research — interpretability, AI-safety / refusal-mechanism study, red-teaming, and robustness evaluation. You assume full responsibility for how you use it and everything it generates; add your own safety and moderation layers before any deployment.

License

Qwen Community License 1.0 (see LICENSE), inherited from the base model Qwen/Qwen3.8-Flash-Next. Abliteration and quantization do not change the underlying license obligations.

Note: per the Qwen Community License, if you operate a Model as a Service or AI Work Assistant business commercially, you must obtain a separate license from Qwen before using this model or its derivatives for commercial purposes. See LICENSE for full terms.