Qwen3.5-0.8B-Heretic — GGUF

GGUF builds of darrellbest/Qwen3.5-0.8B-Heretic. Qwen/Qwen3.5-0.8B with its refusal behaviour removed by Heretic using full-weight Arbitrary-Rank Ablation: 15/100 refusals (original: 98/100) at KL divergence 0.0714. See the main repository for how it was made and measured.

File Quant Size
Qwen3.5-0.8B-Heretic-BF16.gguf BF16 (lossless) 1.56 GB
Qwen3.5-0.8B-Heretic-Q8_0.gguf Q8_0 0.83 GB
Qwen3.5-0.8B-Heretic-Q4_K_M.gguf Q4_K_M 0.54 GB
Qwen3.5-0.8B-Heretic-mmproj-F16.gguf vision projector (F16) 0.20 GB

Converted from the bf16 safetensors with llama.cpp's convert_hf_to_gguf.py and quantized with llama-quantize (no imatrix). The mmproj file carries the vision encoder; load it alongside any of the three for image input.

Checked

Every file was loaded in llama-server with the mmproj. All three answered ordinary prompts correctly and described a test image (a red circle and a blue square) correctly. Thinking-mode reasoning, 4 arithmetic and word problems x 10 seeds with Qwen's recommended sampling: BF16 30/40, Q8_0 30/40, Q4_K_M 27/40 finished and correct (the original model in vLLM: 26/40).

Use

llama-server -m Qwen3.5-0.8B-Heretic-Q8_0.gguf --mmproj Qwen3.5-0.8B-Heretic-mmproj-F16.gguf --jinja -ngl 99

Thinking is on by default; pass "chat_template_kwargs": {"enable_thinking": false} (llama.cpp) or think: false (Ollama) to turn it off per request.

Reduced safety guardrails by design. You are responsible for what you do with it.

The family

Repository Format Size Use it with
Qwen3.5-0.8B-Heretic bf16 safetensors 1.78 GB transformers, vLLM, SGLang
Qwen3.5-0.8B-Heretic-GGUF GGUF BF16 / Q8_0 / Q4_K_M + vision mmproj 1.56 / 0.83 / 0.54 GB + 0.20 GB llama.cpp, Ollama
Qwen3.5-0.8B-Heretic-FP8 FP8 W8A8, compressed-tensors 1.47 GB vLLM
Qwen3.5-0.8B-Heretic-NVFP4 NVFP4, compressed-tensors 1.33 GB vLLM on Blackwell
Downloads last month
363
GGUF
Model size
0.8B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for darrellbest/Qwen3.5-0.8B-Heretic-GGUF

Quantized
(3)
this model