kkxao's picture dbest's picture
Duplicate from darrellbest/Qwen-Image-2.1-PE-I2I-Heretic
d7c7687
|
Raw History Blame Contribute Delete
3.43 kB
metadata
license: other
license_name: qwen-research
license_link: LICENSE
base_model:
  - Qwen/Qwen-Image-2.1-PE-I2I
pipeline_tag: image-text-to-text
tags:
  - heretic
  - abliterated
  - prompt-rewriting
  - image-editing
  - qwen-image

Qwen-Image-2.1-PE-I2I — Heretic (abliterated)

Not affiliated with or endorsed by Alibaba / Qwen. A community derivative of Qwen/Qwen-Image-2.1-PE-I2I, redistributed under the Qwen Research License (copy included as LICENSE). Non-commercial use only; commercial use needs a separate licence from Qwen.

The image-editing prompt rewriter for Qwen-Image-2.1, a fine-tuned Qwen3.5-VL 9B that turns a short edit instruction plus 1–N input images into a detailed English edit prompt, with its refusal behaviour removed by Heretic directional ablation. bf16, same shapes and parameter count as the source; nothing else was changed.

system_prompt.txt is included and required. It defines the output format. It is the unmodified file from the source repo.

Results

Refusals KL divergence
Original 100/100 0 (by definition)
This model (trial 998) 9/100 0.0368

Refusals were measured on mlabonne/harmful_behaviors and KL divergence (damage to ordinary behaviour, first-token distributions) on mlabonne/harmless_alpaca, both Heretic's defaults, with Heretic's default system prompt.

With n = 100 the refusal count is noisy (σ ≈ 3 at this rate), so 6 vs 9 vs 11 are not meaningfully different. Trial 998 was chosen from the Pareto front for having the damage level of the widely used T2I Heretic (pottokao/Qwen-Image-2.1-PE-T2I-Heretic, KL 0.036):

Refusals KL
6/100 0.0431 trial 691
9/100 0.0368 trial 998 (this model)
11/100 0.0330 trial 960
16/100 0.0307 trial 957
20/100 0.0265 trial 975
31/100 0.0194 trial 682

Does it still rewrite edits?

Checked through a real editing app's prompt-enhancer path (the official pe_core output contract, enable_thinking, Qwen's sampling settings) on three edit instructions: a plain one ("make the sky a dramatic sunset") and two borderline ones (horror-film fake blood; a prop handgun on a table). Every answer parsed, and the rewrites were as detailed and specific as the original's.

An honest caveat: the original model did not refuse those image edits either. Like the T2I rewriter, it is relaxed about edgy image subjects and near-total in refusing chat-style harmful instructions, which is what the 100/100 baseline measures. The ablation mainly changes the latter.

Reproduce

  • Heretic 3521f86 (main, 2026-09-21), default config except n_trials = 1000 (200, then 800 more on the same study) and export_strategy = "merge". transformers 5.17.0, torch 2.11.0+cu130, one RTX PRO 6000 Blackwell.
  • Trial 998 parameters: direction_scope = global, direction_index = 19.522, attn.o_proj: max_weight 0.977 at 19.045, min_weight 0.935, min_weight_distance 18.014, mlp.down_proj: max_weight 1.456 at 22.104, min_weight 0.831, min_weight_distance 18.551.

Use

Exactly like the original: load with AutoModelForImageTextToText / AutoProcessor and use system_prompt.txt as the system prompt.