--- license: apache-2.0 base_model: - darrellbest/Qwen3.8-27B-Heretic base_model_relation: quantized pipeline_tag: image-text-to-text library_name: transformers tags: - nvfp4 - compressed-tensors - vllm - heretic - ara - abliterated - qwen3.8 --- # Qwen3.8-27B Heretic NVFP4 NVFP4 (4-bit) build of [darrellbest/Qwen3.8-27B-Heretic](https://huggingface.co/darrellbest/Qwen3.8-27B-Heretic), for **vLLM on NVIDIA Blackwell**, which runs NVFP4 natively. 27 GB instead of 51 GB. ## What is quantized | Part | Precision | |---|---| | Linear layers of the MLPs and the 16 full-attention layers | **NVFP4**, 16-value groups, FP8 scales | | Gated DeltaNet (`linear_attn`) layers, vision tower, embeddings, `lm_head`, MTP | bf16 (unchanged) | This model is 64 layers of which 48 are Gated DeltaNet, and those keep a recurrent state that low precision damages, so they stay in bf16 along with the vision tower. That is why the file is 27 GB rather than roughly 15 GB. Multi-token-prediction weights are preserved in `model-auxiliary.safetensors`, which `save_pretrained` drops for this architecture. Made with [llm-compressor](https://github.com/vllm-project/llm-compressor) 0.13.0 (`scheme="NVFP4"`), calibrated on 64 harmless chat prompts. Format: `compressed-tensors`, `nvfp4-pack-quantized`. ## Use ```sh vllm serve darrellbest/Qwen3.8-27B-Heretic-NVFP4 ``` ## Ablation 0/100 refusals against 98/100 for the original, KL divergence 0.0465. Method and parameters are on the [parent card](https://huggingface.co/darrellbest/Qwen3.8-27B-Heretic). The 4-bit weights were not re-measured for refusals, and the 4-bit activation path is vLLM's, which was not exercised here. > Reduced safety guardrails by design. You are responsible for what you do with it. ## The family | Repository | Format | Size | Use it with | |---|---|---|---| | [Qwen3.8-27B-Heretic](https://huggingface.co/darrellbest/Qwen3.8-27B-Heretic) | bf16 safetensors | 51 GB | transformers, vLLM, SGLang | | [Qwen3.8-27B-Heretic-GGUF](https://huggingface.co/darrellbest/Qwen3.8-27B-Heretic-GGUF) | GGUF BF16 / Q8_0 / Q4_K_M (+ vision) | 51 / 27 / 16 GB | llama.cpp, Ollama | | [Qwen3.8-27B-Heretic-FP8](https://huggingface.co/darrellbest/Qwen3.8-27B-Heretic-FP8) | FP8 W8A8, compressed-tensors | 35 GB | vLLM | | [Qwen3.8-27B-Heretic-NVFP4](https://huggingface.co/darrellbest/Qwen3.8-27B-Heretic-NVFP4) | NVFP4, compressed-tensors | 27 GB | vLLM on Blackwell |