Qwen-Image-2.1-PE-T2I Heretic NVFP4

NVFP4 (4-bit floating point, W4A4) build of darrellbest/Qwen-Image-2.1-PE-T2I-Heretic, the refusal-ablated text-to-image prompt rewriter for Qwen-Image-2.1. For vLLM on NVIDIA Blackwell GPUs, which run NVFP4 natively. 11 GB instead of 18 GB.

Not affiliated with or endorsed by Alibaba / Qwen. Derived from Qwen/Qwen-Image-2.1-PE-T2I under the Qwen Research License (copy included). Non-commercial use only.

system_prompt.txt is included and required, exactly as for the original.

What is quantized

Part Precision
MLP and full-attention projections (32 MLPs, 8 attention layers) NVFP4, 16-value groups, FP8 scales
Gated DeltaNet (linear_attn) layers, vision tower, embeddings, lm_head bf16 (unchanged)

The linear-attention layers carry a recurrent state and the vision tower encodes the input image; both were left in bf16, as other quantizations of this model family do. That is why the file is 11 GB rather than ~6 GB.

Made with llm-compressor 0.13.0 (QuantizationModifier, scheme="NVFP4"), calibrated on 64 samples in the model's real input format: its own system prompt and a short image request. Format: compressed-tensors, nvfp4-pack-quantized.

Checks

Loaded with transformers (which unpacks the 4-bit weights) and run through the prompt-rewriting path: the answer parsed and the rewrite was as detailed and on-target as the bf16 model's. The 4-bit activation path is vLLM's and was not run here, and transformers' emulation is far too slow to use this build outside vLLM.

Use

vllm serve darrellbest/Qwen-Image-2.1-PE-T2I-Heretic-NVFP4

Send the contents of system_prompt.txt as the system message and the image request as the user message, as for the original model.

The family

Model Format Use it with
PE-T2I-Heretic bf16 safetensors transformers / diffusers
PE-T2I-Heretic-GGUF GGUF BF16 / Q8_0 / Q4_K_M llama.cpp, Ollama
PE-T2I-Heretic-NVFP4 NVFP4 (compressed-tensors) vLLM on Blackwell
PE-I2I-Heretic bf16 safetensors transformers / diffusers
PE-I2I-Heretic-GGUF GGUF BF16 / Q8_0 / Q4_K_M (+ mmproj) llama.cpp, Ollama
PE-I2I-Heretic-NVFP4 NVFP4 (compressed-tensors) vLLM on Blackwell

T2I rewrites a short request into a detailed prompt for new images; I2I rewrites an edit instruction, reading the image being edited. Both are prompt rewriters for Qwen-Image-2.1, not image generators.

Downloads last month
65
Safetensors
Model size
9B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for darrellbest/Qwen-Image-2.1-PE-T2I-Heretic-NVFP4

Quantized
(2)
this model