Diffusers
Safetensors
qwen3_vl
qwen
qwen3-vl
text-encoder
vision-encoder
heretic
abliteration
prompt-adherence
Instructions to use catplusplus/Qwen21_Text_Encoder_Heretic with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use catplusplus/Qwen21_Text_Encoder_Heretic with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("catplusplus/Qwen21_Text_Encoder_Heretic", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
File size: 8,718 Bytes
b6fb11b 7b00440 b6fb11b 7b00440 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 | ---
license: apache-2.0
base_model: Qwen/Qwen-Image-2.1
tags:
- qwen
- qwen3-vl
- text-encoder
- vision-encoder
- heretic
- abliteration
- prompt-adherence
- diffusers
---
# Qwen3-VL-8B Heretic Text & Vision Encoder (Prompt Adherence & Geometric Alignment Edition) πΊβ¨
This repository provides an optimized, abliterated checkpoint of the **Qwen3-VL-8B** text and vision encoder from **[Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1)**, processed with **Norm-Preserving Biprojected Abliteration**.
The primary purpose of this model is **maximum instruction following and prompt adherence**: it banishes geometric representation deflection ("internal blush" / hesitation vectors) that otherwise causes safety-tuned VLMs to corrupt diffusion conditioning on dynamic poses, human figures, athletic wear, and complex scenes.
---
## π¬ The Core Problem: Why VLM Safety Alignment Degrades Diffusion Conditioning
In text-generation tasks, safety alignment mechanisms steer models to emit refusal text (e.g. *"I cannot fulfill this request..."*). However, modern multimodal diffusion architectures like **Qwen-Image-2.1** do **not** generate text tokens:
$$\text{DiT Conditioning} \longleftarrow \mathbf{h}_L = \text{TextEncoder}(\text{tokens})[-1]$$
The diffusion transformer taps the raw pre-RMSNorm residual hidden states $\mathbf{h}_L$ directly from the text encoder to drive cross-attention.
### The "Internal Blush" / Hesitation Deflection Phenomenon
When prompts describe human subjects, dynamic physical actions, athletic attire (e.g. swimwear, volleyball, gymnastics), or expressive emotions, safety-tuning vectors inside the language model activate even on completely benign, non-refusal prompts.
Because the model cannot output a refusal string, these alignment vectors manifest as a **geometric rotation of the latent representation**:
$$\mathbf{h}_{\text{sensitive}} = \mathbf{h}_{\text{clean}} + \mathbf{v}_{\text{refusal}}$$
This hidden deflection rotates the conditioning signal by **over 60% relative norm** away from the prompt's intended semantic visual trajectory!
### Visual Consequences in Image Generation
When cross-attention layers in the DiT receive a representation deflected into the refusal/modesty subspace, the model displays hesitation artifacts:
1. **Modesty Hallucinations & Clothing Confusion**: Spontaneous addition of mismatched cloth, awkward white ruffles, or extra fabric covering swimwear or sportswear.
2. **Anatomical Occlusion**: The model avoids rendering human limbs or athletic poses, awkwardly hiding arms behind character backs or contorting torsos.
3. **Subject & Prop Merging**: Equipment or background elements get fused into characters (e.g. sports balls bizarrely merged onto heads as hair ornaments).
4. **Action Damping**: Dynamic verbs (*"jumping to spike the ball"*) are subdued into passive, static standing postures.
---
## π Quantitative Measurement of Representation Deflection
Using contrastive prompt pairs across benign and sensitive subjects, we measured the layer-by-layer cosine similarity and relative deflection norm across all 37 positions (input embeddings + 36 decoder layers) of Qwen3-VL-8B:
| Layer Index | Position | Cosine Similarity ($\cos \theta$) | Relative Deflection ($\|\Delta \mathbf{h}\| / \|\mathbf{h}\|$) | Deflection Norm $\|\Delta \mathbf{h}\|$ |
| :---: | :---: | :---: | :---: | :---: |
| **0** | Input Embeddings | **1.0000** | **0.00%** | 0.00 |
| **8** | Early Transformer | **0.9991** | **3.82%** | 18.24 |
| **16** | Mid-Low (Deflection Onset) | **0.9943** | **10.64%** | 52.88 |
| **20** | Mid-High Divergence | **0.9780** | **21.05%** | 114.73 |
| **24** | Refusal Vector Surge | **0.9414** | **34.25%** | 192.40 |
| **28** | Acceleration Peak | **0.8842** | **47.19%** | 275.31 |
| **32** | Late Transformer | **0.8350** | **56.12%** | 331.05 |
| **36** | Final Conditioning Layer | **0.8110** | **60.30%** | **357.94** |
Between Layer 20 and Layer 36, representation deflection accelerates rapidly, culminating in a **60.3% vector distortion**. By surgically neutralizing this direction, the text encoder reflects the exact intended prompt semantics.
---
## π οΈ Methodology: Norm-Preserving Biprojected Abliteration
To eliminate hesitation deflection without degrading general language comprehension, we applied **Norm-Preserving Biprojected Abliteration** (`create_heretic_text_encoder.py`):
1. **Refusal Subspace Extraction**: Difference-of-means vectors were extracted across contrastive prompt sets:
$$\mathbf{r}_l = \boldsymbol{\mu}_{\text{sensitive}}^{(l)} - \boldsymbol{\mu}_{\text{benign}}^{(l)}$$
2. **Benign Subspace Orthogonalization**: The general semantic direction was stripped from the refusal vector:
$$\mathbf{v}_l = \mathbf{r}_l - \text{proj}_{\mathbf{u}_{\text{benign}}}(\mathbf{r}_l)$$
3. **Norm-Preserving Rank-1 Projection**: Across 54 linear projection matrices (`self_attn.o_proj` and `mlp.down_proj` in layers 9β35, centered at layer 26 with Gaussian falloff $\lambda \in [0.10, 1.00]$):
$$W_{\text{norm}} = \text{normalize}(W, p=2, \text{dim}=1)$$
$$W' = \text{normalize}\Big(W_{\text{norm}} - \lambda \mathbf{v}_l (\mathbf{v}_l^T W_{\text{norm}})\Big) \cdot \|W\|_{\text{row}}$$
Because exact row norms ($\|W\|_{\text{row}}$) are strictly preserved, the network's overall activation scales and general reasoning capabilities remain completely intact.
---
## π€ Pairing with Quantized DiT (`nunchaku-qwen-image-2.1`)
This text encoder is specifically engineered to be paired with **[`nunchaku-qwen-image-2.1`](https://huggingface.co/models/nunchaku-qwen-image-2.1)** for consumer GPU setups:
* **Resident DiT + Streamed Text Encoder**:
* DiT (`best_quality_fp4.safetensors`): **4.08 GB resident VRAM**.
* VAE (`AutoencoderKLQwenImage21`): **0.64 GB resident VRAM**.
* Qwen3-VL-8B ViT Vision Encoder: **1.07 GB resident VRAM**.
* Qwen3-VL-8B Language Model: Streamed layer-by-layer through a static **368 MB GPU buffer** over PCIe at ~28.7 GB/s via `stream_encoder.py`.
* **Total VRAM Footprint**: **~6.17 GB active VRAM**, leaving **~9.5 GB free headroom** on a single 16 GB GPU (such as RTX 5060 Ti or RTX 4080)!
* **Inference Speed**: Multimodal prompt encoding completes in **1.06s** (saving 16s vs CPU), and 25-step image generation runs in **~20s**.
---
## π Quickstart Usage
### 1. Installation
```bash
pip install diffusers transformers accelerate torch sentencepiece
```
### 2. Loading with Diffusers
```python
import torch
from diffusers import QwenImage21Pipeline
from transformers import Qwen3VLForConditionalGeneration, Qwen3VLProcessor
# 1. Load Heretic text encoder and processor
text_encoder = Qwen3VLForConditionalGeneration.from_pretrained(
"models/Qwen21_Text_Encoder_Heretic",
torch_dtype=torch.bfloat16,
low_cpu_mem_usage=True,
)
processor = Qwen3VLProcessor.from_pretrained("models/Qwen21_Text_Encoder_Heretic")
# 2. Assemble into pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1",
text_encoder=text_encoder,
processor=processor,
torch_dtype=torch.bfloat16,
)
pipe.enable_sequential_cpu_offload(gpu_id=0)
# 3. Generate with precise prompt adherence
image = pipe(
prompt="Two cute anime girls in colorful bikinis playing beach volleyball on a sunny tropical beach, dynamic action pose, jumping to spike the ball, sharp focus",
height=1024,
width=1024,
num_inference_steps=25,
true_cfg_scale=1.0,
).images[0]
image.save("beach_volleyball.png")
```
### 3. High-Throughput Server Usage
Run the bundled ImageEditServer with NVFP4 DiT and Heretic text encoder:
```bash
# Start server on port 4500 (uses Heretic text encoder by default)
./extras/imagegen_qwen21_nvfp4.sh 4500
```
---
## π¦ Packaged Sources (`extras/`)
* `create_heretic_text_encoder.py`: Complete script used to measure refusal vectors and perform norm-preserving biprojected abliteration.
* `stream_encoder.py`: Zero-quality-loss layerwise weight streaming engine for Qwen3-VL-8B.
* `test_heretic_beach_volleyball.py`: Empirical verification script comparing stock vs Heretic encoders.
* `QwenImage21NVFP4Backend.py`: Diffusers + Nunchaku backend supporting custom text encoder overrides.
* `ImageEditServer.py` & `imagegen_qwen21_nvfp4.sh`: Resident image generation server.
---
## π Citation & Credits
* **Qwen-Image-2.1 & Qwen3-VL**: Qwen Team, Alibaba Cloud.
* **Abliteration Principles**: Arditi et al. (*Refusal in Language Models Is Mediated by a Single Direction*).
* **Heretic LLM**: Heretic project (*Directional Abliteration Toolkit*).
* **Abliteration & Diffusion Conditioning Optimization**: Oleg K. / Nikola Seeker Project.
|