DeepSeek-R1-Qwen3-8B-heretic

RACER IS OP

A decensored variant of deepseek-ai/DeepSeek-R1-0528-Qwen3-8B, produced with Heretic v1.4.0 (directional ablation / "abliteration"). DeepSeek's R1-0528 distill of Qwen3-8B keeps its native <think> reasoning traces; refusal behaviour is suppressed here via targeted weight edits to the attention output and MLP down-projections rather than fine-tuning, so the R1 reasoning tuning is left largely intact.

Who this is for: developers who want uncensored R1-style reasoning at 8B - visible <think> traces, refusals down from 87/100 to 3/100 at a KL of just 0.0082, one of the tightest edits in this collection. At 8B with 128K context it runs on a single consumer GPU via GGUF.

Runs on your gaming PC

Full GGUF ladder included - pick the quant that fits your card:

Your GPU Recommended quant Weights
RTX 4090 / 5090 (24 GB) Q8_0 8.11 GB
RTX 4080 / 5080 / 4060 Ti 16G (16 GB) Q6_K 6.26 GB
RTX 3060 / 4070 / 5070 (12 GB) Q5_K_M 5.45 GB
RTX 4060 / 3070 (8 GB) Q4_K_M 4.68 GB
CPU-only / Apple Silicon Q4_K_M 4.68 GB

Weights only, at this model's native 8.2B size; add ~1 GB per 32K of context. OOM? Drop one quant level. Headroom to spare? Go one up.

Abliteration parameters

Trial 98 of a 200-trial Heretic run (seed 750233003).

Parameter Value
direction_index 19.81
attn.o_proj.max_weight 1.44
attn.o_proj.max_weight_position 21.05
attn.o_proj.min_weight 1.38
attn.o_proj.min_weight_distance 20.94
mlp.down_proj.max_weight 1.43
mlp.down_proj.max_weight_position 25.81
mlp.down_proj.min_weight 1.19
mlp.down_proj.min_weight_distance 20.76

Performance

Metric This model Original model (deepseek-ai/DeepSeek-R1-0528-Qwen3-8B)
KL divergence 0.0082 0 (by definition)
Refusals 3/100 87/100

Refusals on the harmful evaluation set drop from 87/100 to 3/100 - effectively total removal. KL divergence of 0.0082 is one of the tightest edits in this collection, so the R1 reasoning style and formatting stay very close to the original. direction_index is a single index (19.81) rather than per-layer, which is why the refusal removal lands this completely without needing a wider edit.

Why abliteration instead of fine-tuning

Fine-tuning a "helpful" persona on top of RLHF'd refusals fights the base model's training and tends to degrade coherence. Abliteration instead finds and edits the specific weight directions responsible for refusal, leaving the rest of the network (and its capabilities) untouched. See the Heretic repo and the original abliteration writeup for the mechanism.

Made with ❤️ by RACER IS OP — follow for more uncensored models

Files

Safetensors

File Size
model-00001-of-00004.safetensors 4.57 GB
model-00002-of-00004.safetensors 4.58 GB
model-00003-of-00004.safetensors 4.63 GB
model-00004-of-00004.safetensors 1.48 GB

BF16, ~8.2B parameters. The reproduce/ directory carries the full Heretic recipe - config.toml, requirements.txt, the Optuna study journal, and SHA-256 sums - so this exact model can be regenerated bit-for-bit. Reproduce it with heretic --reproduce reproduce/reproduce.json.

GGUF quantizations

Full quantization set (F16 + Q4_K_M, Q5_K_M, Q6_K, Q8_0) produced with llama.cpp.

File Format Size
DeepSeek-R1-Qwen3-8B-heretic-F16.gguf GGUF F16 15.26 GB
DeepSeek-R1-Qwen3-8B-heretic-Q4_K_M.gguf GGUF Q4_K_M 4.68 GB
DeepSeek-R1-Qwen3-8B-heretic-Q5_K_M.gguf GGUF Q5_K_M 5.45 GB
DeepSeek-R1-Qwen3-8B-heretic-Q6_K.gguf GGUF Q6_K 6.26 GB
DeepSeek-R1-Qwen3-8B-heretic-Q8_0.gguf GGUF Q8_0 8.11 GB

Qwen3 architecture (qwen3) - loads natively in llama.cpp / Ollama / LM Studio / Jan.

Tokenizer note: the repo's tokenizer.json was re-saved by Heretic with a Metaspace pre-tokenizer wrapper and the snapshot declares the slow LlamaTokenizer, a combination the GGUF converter rejects. The GGUFs were built with the base model's ByteLevel pre-tokenizer and Qwen2Tokenizer (vocab, merges, and all 28 added tokens verified identical; token IDs verified identical to reference Qwen3 tokenization), so the embedded qwen2 pre-tokenizer behaves exactly as the base model does.

Run llama serve -hf saidutta69/DeepSeek-R1-Qwen3-8B-heretic to pull the default quant.

Quickstart

# llama.cpp
llama serve -hf saidutta69/DeepSeek-R1-Qwen3-8B-heretic
# transformers
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "saidutta69/DeepSeek-R1-Qwen3-8B-heretic"
model = AutoModelForCausalLM.from_pretrained(model_name, dtype="auto", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)

messages = [{"role": "user", "content": "Prove that the square root of 2 is irrational. Show each step of the reasoning."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=True,
                                        return_dict=True, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=2048)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

Also runnable via Ollama, LM Studio, Jan, vLLM, SGLang.

Thinking traces

This is the R1 reasoning edition - the model emits a <think>...</think> block before its answer. For latency-sensitive agent loops where you want tokens directly, use a non-thinking model from this collection instead; for hard problems that benefit from visible reasoning, this one is the right pick.

Model details

Architecture Qwen3ForCausalLM (decoder-only dense transformer)
Parameters ~8.2B
Layers / heads 36 layers, 32 attention heads, 8 KV heads (GQA), head dim 128
Hidden / intermediate 4096 / 12288 (SwiGLU)
Position embedding RoPE (YaRN, factor 4.0), theta = 1,000,000
Context length 131,072
Vocab 151,936
Precision bfloat16
Base model deepseek-ai/DeepSeek-R1-0528-Qwen3-8B

Responsible use

Refusal suppression is deliberate and works as intended: this model will comply with requests the base model would refuse, including some it shouldn't. There is no safety filtering layered on top. You are responsible for how you deploy it — don't put this behind an unmoderated public-facing endpoint serving third parties. It inherits R1/Qwen3's factual limitations and biases; abliteration removes refusal directions, it doesn't add capability or judgment.

License

Inherits the mit license from the base model.

Related

Downloads last month
546
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for saidutta69/DeepSeek-R1-Qwen3-8B-heretic

Finetuned
(67)
this model

Collection including saidutta69/DeepSeek-R1-Qwen3-8B-heretic