Qwen3.5-0.8B-Heretic NVFP4
NVFP4 build of darrellbest/Qwen3.5-0.8B-Heretic for vLLM on NVIDIA Blackwell, which runs NVFP4 natively. 1.33 GB instead of 1.78 GB. Qwen/Qwen3.5-0.8B with its refusal behaviour removed by Heretic using full-weight Arbitrary-Rank Ablation: 15/100 refusals (original: 98/100) at KL divergence 0.0714. See the main repository for how it was made and measured.
What is quantized
| Part | Precision |
|---|---|
| MLP linear layers, and the attention projections of the full-attention layers | NVFP4, 16-value groups, FP8 scales |
| Everything else | unchanged |
Unchanged: the vision encoder, the Gated DeltaNet (linear_attn) layers, the multi-token-prediction block, the embeddings (tied to lm_head) and norms stay bf16 (the DeltaNet A_log/norm parameters float32, as in the original); the recurrent DeltaNet state is sensitive to low precision, and the 248k-token embedding table is a large share of a model this size. Made with llm-compressor 0.13.0
(scheme="NVFP4", calibrated on 64 harmless chat prompts). The multi-token-prediction
weights, which the quantized save drops, were copied back unchanged into model-auxiliary.safetensors.
Checked
Loaded in vLLM 0.30.0 on an RTX PRO 6000 Blackwell: ordinary prompts answered correctly and the test image (a red circle and a blue square) described correctly. Thinking-mode reasoning, 4 arithmetic and word problems x 10 seeds: 24/40 finished and correct (bf16 Heretic 27/40, original 26/40). Throughput: ~380 tok/s single-stream, ~6,400 tok/s aggregate at batch 32 (512-token generations).
Use
vllm serve darrellbest/Qwen3.5-0.8B-Heretic-NVFP4
The NVFP4 weights were not re-measured for refusals.
Reduced safety guardrails by design. You are responsible for what you do with it.
The family
| Repository | Format | Size | Use it with |
|---|---|---|---|
| Qwen3.5-0.8B-Heretic | bf16 safetensors | 1.78 GB | transformers, vLLM, SGLang |
| Qwen3.5-0.8B-Heretic-GGUF | GGUF BF16 / Q8_0 / Q4_K_M + vision mmproj | 1.56 / 0.83 / 0.54 GB + 0.20 GB | llama.cpp, Ollama |
| Qwen3.5-0.8B-Heretic-FP8 | FP8 W8A8, compressed-tensors | 1.47 GB | vLLM |
| Qwen3.5-0.8B-Heretic-NVFP4 | NVFP4, compressed-tensors | 1.33 GB | vLLM on Blackwell |
- Downloads last month
- 20
Model tree for darrellbest/Qwen3.5-0.8B-Heretic-NVFP4
Base model
Qwen/Qwen3.5-0.8B-Base