How to use from
Docker Model Runner
docker model run hf.co/darrellbest/Qwen3.8-27B-Heretic-NVFP4
Quick Links

Qwen3.8-27B Heretic NVFP4

NVFP4 (4-bit) build of darrellbest/Qwen3.8-27B-Heretic, for vLLM on NVIDIA Blackwell, which runs NVFP4 natively. 27 GB instead of 51 GB.

What is quantized

Part Precision
Linear layers of the MLPs and the 16 full-attention layers NVFP4, 16-value groups, FP8 scales
Gated DeltaNet (linear_attn) layers, vision tower, embeddings, lm_head, MTP bf16 (unchanged)

This model is 64 layers of which 48 are Gated DeltaNet, and those keep a recurrent state that low precision damages, so they stay in bf16 along with the vision tower. That is why the file is 27 GB rather than roughly 15 GB. Multi-token-prediction weights are preserved in model-auxiliary.safetensors, which save_pretrained drops for this architecture.

Made with llm-compressor 0.13.0 (scheme="NVFP4"), calibrated on 64 harmless chat prompts. Format: compressed-tensors, nvfp4-pack-quantized.

Use

vllm serve darrellbest/Qwen3.8-27B-Heretic-NVFP4

Ablation

0/100 refusals against 98/100 for the original, KL divergence 0.0465. Method and parameters are on the parent card. The 4-bit weights were not re-measured for refusals, and the 4-bit activation path is vLLM's, which was not exercised here.

Reduced safety guardrails by design. You are responsible for what you do with it.

The family

Repository Format Size Use it with
Qwen3.8-27B-Heretic bf16 safetensors 51 GB transformers, vLLM, SGLang
Qwen3.8-27B-Heretic-GGUF GGUF BF16 / Q8_0 / Q4_K_M (+ vision) 51 / 27 / 16 GB llama.cpp, Ollama
Qwen3.8-27B-Heretic-FP8 FP8 W8A8, compressed-tensors 35 GB vLLM
Qwen3.8-27B-Heretic-NVFP4 NVFP4, compressed-tensors 27 GB vLLM on Blackwell
Downloads last month
35
Safetensors
Model size
19B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for darrellbest/Qwen3.8-27B-Heretic-NVFP4

Base model

Qwen/Qwen3.8-27B
Quantized
(3)
this model