๐Ÿฑ All-GGUF Qwen-Image-2.1: text encoder GGUF + DiT GGUF (Q8_0 / Q6_K / Q4_K_M, loads in stock ComfyUI-GGUF) + PE-T2I GGUF.

The 4 model-0000x-of-00004 shards are the HF transformers checkpoint (bf16) โ€” they do not load in ComfyUI. ComfyUI text encoders use a different key layout (the model.language_model. prefix is dropped when repacking), so pointing ComfyUI at these shards will not work.

For ComfyUI use one of these instead โ€” all keep the vision tower, which 2.1 needs for editing. Mac / non-CUDA: use the single file qwen3vl_8b_bf16_heretic.safetensors in this repo (full bf16, 17.5 GB โ€” also in the GGUF repo) โ€” put it in models/text_encoders/ and load it with the stock CLIPLoader, type qwen_image, then TextEncodeQwenImage21. It is the only format here that needs no CUDA-specific kernels.

Repo File Loader
this repo qwen3vl_8b_bf16_heretic.safetensors (bf16 single file, any device) CLIPLoader, type qwen_image
โ€ฆ-NVFP4 qwen3vl_8b_nvfp4_heretic.safetensors CLIPLoader, type qwen_image
โ€ฆ-W4A8 qwen3vl_8b_w4a8_heretic.safetensors CLIPLoader, type qwen_image
โ€ฆ-GGUF qwen3vl_8b_heretic-Q4_K_M.gguf + mmproj-โ€ฆ-f16.gguf CLIPLoaderGGUF + the ComfyUI-GGUF-Qwen3VL-TE add-on node (without it: [1, 512, 12288] shape error)

You also need a ComfyUI new enough to know QwenImage21: 0.34.2 does not, 0.36.0 does. If TextEncodeQwenImage21 is missing from your node list, that is why.

Use this repo for transformers / diffusers / vLLM, or as the base for your own quantization.

Qwen-Image-2.1 Text Encoder โ€” Heretic (Abliterated)

Not affiliated with, or endorsed by, Alibaba / Qwen. Community derivative (refusal-ablated) of Qwen/Qwen3-VL-8B-Instruct โ€” the model Qwen-Image-2.1 uses, unmodified, as its text encoder. Qwen releases that model under Apache-2.0, so this derivative is redistributed under Apache-2.0 (see LICENSE and NOTICE).

The text encoder of Qwen/Qwen-Image-2.1 (a Qwen3-VL-8B-Instruct) with refusal behaviour removed via Heretic directional ablation.

Drop-in replacement for the stock text encoder. Weights are bf16, same shapes, same parameter count โ€” nothing else was changed.

Results

Refusals KL divergence
Original (measured baseline) 100/100 0 (by definition)
This model 5/100 0.0220

Measured by Heretic on mlabonne/harmful_behaviors (test split) for refusals and mlabonne/harmless_alpaca for KL divergence โ€” i.e. lower refusals and lower distribution shift on benign inputs.

Independent verification

Refusal rate and general capability were re-checked with a separate script (different code, different refusal keyword set) rather than trusting the optimiser's own numbers:

  • Refusals: 0/20 on held-out harmful prompts
  • Benign questions: 4/4 correct and coherent (e.g. "What is the capital of France?" โ†’ *"The capital of France is Parisโ€ฆ"*)

โ˜… Search budget matters โ€” measured, not assumed

Heretic's documented defaults are n_trials = 200, n_startup_trials = 60. A first run with 100/20 was done for comparison:

Run trials / startup Best balanced result
v1 100 / 20 9/100 @ KL 0.0338
v2 (this model) 200 / 60 5/100 @ KL 0.0220

Doubling the search budget nearly halved the refusal rate and cut KL divergence by a third. 100 trials is not enough for this model. If you are abliterating something similar, use the documented defaults.

Pareto front

The optimiser returns a Pareto front; this release uses the knee point, not the extreme:

index refusals KL note
0 4/100 0.0859 1 fewer refusal costs 3.9ร— the KL
1 5/100 0.0220 โ† released
2 28/100 0.0165 0.0055 less KL costs +23pp refusals

Reproduction

uvx --from "git+https://github.com/p-e-w/heretic@3521f8648a0dccf6e12a92666862632235fac7e6" heretic \
  --model <path to Qwen-Image-2.1/text_encoder + processor, flattened> \
  --dtypes bfloat16 --device-map auto \
  --max-memory '{"0":"14GiB","1":"14GiB"}' \
  --offload-outputs-to-cpu --max-batch-size 32 \
  --n-trials 200 --n-startup-trials 60 \
  --study-checkpoint-dir <ckpt> \
  --trial-index 1 --model-action save \
  --save-directory <out> --export-strategy MERGE

Hardware: 2ร— RTX 5070 Ti (16 GB each), 48 min for 200 trials (14.5 s/trial).

Gotchas worth knowing

  • Pin the commit. git+โ€ฆ/heretic without a revision is a moving target; the commit above reports v2.0.0.dev0. The PyPI release heretic-llm==1.4.0 is older and rejects --trial-index / --model-action / --save-directory.
  • --trial-index is the index into the sorted Pareto front, not the Optuna trial id. Passing a trial id silently falls back to the interactive menu.
  • --checkpoint-action continue replaces the entire settings object with the one stored in the checkpoint (main.py:404-407), discarding your CLI flags. To export a different trial afterwards you must patch the settings stored in the study journal, not the command line.
  • Finishing a run opens an interactive TUI; with stdin=/dev/null it raises EOFError. Pass a valid --trial-index to avoid it.

Usage

Standard transformers:

from transformers.models.qwen3_vl import Qwen3VLForConditionalGeneration
model = Qwen3VLForConditionalGeneration.from_pretrained(
    "pottokao/Qwen-Image-2.1-Text-Encoder-Heretic", dtype="bfloat16", device_map="auto")

For ComfyUI, note that Comfy-Org's repack strips the model.language_model. prefix (model.language_model.layers.N.โ€ฆ โ†’ model.layers.N.โ€ฆ). Weights quantized straight from this HF layout will not load in ComfyUI until the keys are remapped.

Notes

  • Only the text encoder is modified. The DiT and VAE of Qwen-Image-2.1 are untouched.
  • Ablation targets o_proj and down_proj (Heretic's defaults for this model).
  • Quantizing this model behaves the same as quantizing the original: NVFP4 round-trip error measured 9.52 % on ablated layers vs 9.51 % on untouched layers vs 9.44 % on the stock encoder โ€” ablation does not make the weights harder to quantize, so the same recipe applies.
Downloads last month
5,918
Safetensors
Model size
9B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for pottokao/Qwen-Image-2.1-Text-Encoder-Heretic

Finetuned
(614)
this model
Adapters
1 model
Finetunes
6 models
Quantizations
5 models

Space using pottokao/Qwen-Image-2.1-Text-Encoder-Heretic 1