Qwen3.5-9B-Heretic NVFP4

NVFP4 build of darrellbest/Qwen3.5-9B-Heretic for vLLM on NVIDIA Blackwell, which runs NVFP4 natively. 11.72 GB instead of 19.34 GB. Qwen/Qwen3.5-9B with its refusal behaviour removed by Heretic using full-weight Arbitrary-Rank Ablation: 5/100 refusals (original: 100/100) at KL divergence 0.0403. See the main repository for how it was made and measured.

What is quantized

Part Precision
MLP linear layers, and the attention projections of the full-attention layers NVFP4, 16-value groups, FP8 scales
Everything else unchanged

Unchanged: the vision encoder, the Gated DeltaNet (linear_attn) layers, the multi-token-prediction block, the embeddings (tied to lm_head) and norms stay bf16 (the DeltaNet A_log/norm parameters float32, as in the original); the recurrent DeltaNet state is sensitive to low precision, and the 248k-token embedding table is a large share of a model this size. Made with llm-compressor 0.13.0 (scheme="NVFP4", calibrated on 64 harmless chat prompts). The multi-token-prediction weights, which the quantized save drops, were copied back unchanged into model-auxiliary.safetensors.

Checked

Loaded in vLLM 0.30.0 on an RTX PRO 6000 Blackwell: ordinary prompts answered correctly and the test image (a red circle and a blue square) described correctly. Thinking-mode reasoning, 4 arithmetic and word problems x 10 seeds with Qwen's recommended sampling: 34/40 finished and correct (bf16 Heretic 40/40, original 40/40); all six misses ran out of the 4,000-token budget while still thinking, none answered wrongly, so at 4 bits it reasons at more length rather than less accurately. Throughput: ~100 tok/s single-stream, ~2,550 tok/s aggregate at batch 32 (512-token generations; bf16: ~57 / ~1,550).

Use

vllm serve darrellbest/Qwen3.5-9B-Heretic-NVFP4

The NVFP4 weights were not re-measured for refusals.

Reduced safety guardrails by design. You are responsible for what you do with it.

The family

Repository Format Size Use it with
Qwen3.5-9B-Heretic bf16 safetensors 19.34 GB transformers, vLLM, SGLang
Qwen3.5-9B-Heretic-GGUF GGUF BF16 / Q8_0 / Q4_K_M + vision mmproj 18.41 / 9.79 / 5.78 GB + 0.92 GB llama.cpp, Ollama
Qwen3.5-9B-Heretic-FP8 FP8 W8A8, compressed-tensors 14.04 GB vLLM
Qwen3.5-9B-Heretic-NVFP4 NVFP4, compressed-tensors 11.72 GB vLLM on Blackwell
Downloads last month
24
Safetensors
Model size
10B params
Tensor type
F32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for darrellbest/Qwen3.5-9B-Heretic-NVFP4

Finetuned
Qwen/Qwen3.5-9B
Quantized
(3)
this model