File size: 8,718 Bytes
b6fb11b
 
7b00440
 
 
 
 
 
 
 
 
 
b6fb11b
7b00440
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
---
license: apache-2.0
base_model: Qwen/Qwen-Image-2.1
tags:
- qwen
- qwen3-vl
- text-encoder
- vision-encoder
- heretic
- abliteration
- prompt-adherence
- diffusers
---

# Qwen3-VL-8B Heretic Text & Vision Encoder (Prompt Adherence & Geometric Alignment Edition) 🌺✨

This repository provides an optimized, abliterated checkpoint of the **Qwen3-VL-8B** text and vision encoder from **[Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1)**, processed with **Norm-Preserving Biprojected Abliteration**.

The primary purpose of this model is **maximum instruction following and prompt adherence**: it banishes geometric representation deflection ("internal blush" / hesitation vectors) that otherwise causes safety-tuned VLMs to corrupt diffusion conditioning on dynamic poses, human figures, athletic wear, and complex scenes.

---

## πŸ”¬ The Core Problem: Why VLM Safety Alignment Degrades Diffusion Conditioning

In text-generation tasks, safety alignment mechanisms steer models to emit refusal text (e.g. *"I cannot fulfill this request..."*). However, modern multimodal diffusion architectures like **Qwen-Image-2.1** do **not** generate text tokens:

$$\text{DiT Conditioning} \longleftarrow \mathbf{h}_L = \text{TextEncoder}(\text{tokens})[-1]$$

The diffusion transformer taps the raw pre-RMSNorm residual hidden states $\mathbf{h}_L$ directly from the text encoder to drive cross-attention.

### The "Internal Blush" / Hesitation Deflection Phenomenon

When prompts describe human subjects, dynamic physical actions, athletic attire (e.g. swimwear, volleyball, gymnastics), or expressive emotions, safety-tuning vectors inside the language model activate even on completely benign, non-refusal prompts. 

Because the model cannot output a refusal string, these alignment vectors manifest as a **geometric rotation of the latent representation**:

$$\mathbf{h}_{\text{sensitive}} = \mathbf{h}_{\text{clean}} + \mathbf{v}_{\text{refusal}}$$

This hidden deflection rotates the conditioning signal by **over 60% relative norm** away from the prompt's intended semantic visual trajectory!

### Visual Consequences in Image Generation

When cross-attention layers in the DiT receive a representation deflected into the refusal/modesty subspace, the model displays hesitation artifacts:
1. **Modesty Hallucinations & Clothing Confusion**: Spontaneous addition of mismatched cloth, awkward white ruffles, or extra fabric covering swimwear or sportswear.
2. **Anatomical Occlusion**: The model avoids rendering human limbs or athletic poses, awkwardly hiding arms behind character backs or contorting torsos.
3. **Subject & Prop Merging**: Equipment or background elements get fused into characters (e.g. sports balls bizarrely merged onto heads as hair ornaments).
4. **Action Damping**: Dynamic verbs (*"jumping to spike the ball"*) are subdued into passive, static standing postures.

---

## πŸ“Š Quantitative Measurement of Representation Deflection

Using contrastive prompt pairs across benign and sensitive subjects, we measured the layer-by-layer cosine similarity and relative deflection norm across all 37 positions (input embeddings + 36 decoder layers) of Qwen3-VL-8B:

| Layer Index | Position | Cosine Similarity ($\cos \theta$) | Relative Deflection ($\|\Delta \mathbf{h}\| / \|\mathbf{h}\|$) | Deflection Norm $\|\Delta \mathbf{h}\|$ |
| :---: | :---: | :---: | :---: | :---: |
| **0** | Input Embeddings | **1.0000** | **0.00%** | 0.00 |
| **8** | Early Transformer | **0.9991** | **3.82%** | 18.24 |
| **16** | Mid-Low (Deflection Onset) | **0.9943** | **10.64%** | 52.88 |
| **20** | Mid-High Divergence | **0.9780** | **21.05%** | 114.73 |
| **24** | Refusal Vector Surge | **0.9414** | **34.25%** | 192.40 |
| **28** | Acceleration Peak | **0.8842** | **47.19%** | 275.31 |
| **32** | Late Transformer | **0.8350** | **56.12%** | 331.05 |
| **36** | Final Conditioning Layer | **0.8110** | **60.30%** | **357.94** |

Between Layer 20 and Layer 36, representation deflection accelerates rapidly, culminating in a **60.3% vector distortion**. By surgically neutralizing this direction, the text encoder reflects the exact intended prompt semantics.

---

## πŸ› οΈ Methodology: Norm-Preserving Biprojected Abliteration

To eliminate hesitation deflection without degrading general language comprehension, we applied **Norm-Preserving Biprojected Abliteration** (`create_heretic_text_encoder.py`):

1. **Refusal Subspace Extraction**: Difference-of-means vectors were extracted across contrastive prompt sets:
   $$\mathbf{r}_l = \boldsymbol{\mu}_{\text{sensitive}}^{(l)} - \boldsymbol{\mu}_{\text{benign}}^{(l)}$$
2. **Benign Subspace Orthogonalization**: The general semantic direction was stripped from the refusal vector:
   $$\mathbf{v}_l = \mathbf{r}_l - \text{proj}_{\mathbf{u}_{\text{benign}}}(\mathbf{r}_l)$$
3. **Norm-Preserving Rank-1 Projection**: Across 54 linear projection matrices (`self_attn.o_proj` and `mlp.down_proj` in layers 9–35, centered at layer 26 with Gaussian falloff $\lambda \in [0.10, 1.00]$):
   $$W_{\text{norm}} = \text{normalize}(W, p=2, \text{dim}=1)$$
   $$W' = \text{normalize}\Big(W_{\text{norm}} - \lambda \mathbf{v}_l (\mathbf{v}_l^T W_{\text{norm}})\Big) \cdot \|W\|_{\text{row}}$$

Because exact row norms ($\|W\|_{\text{row}}$) are strictly preserved, the network's overall activation scales and general reasoning capabilities remain completely intact.

---

## 🀝 Pairing with Quantized DiT (`nunchaku-qwen-image-2.1`)

This text encoder is specifically engineered to be paired with **[`nunchaku-qwen-image-2.1`](https://huggingface.co/models/nunchaku-qwen-image-2.1)** for consumer GPU setups:

* **Resident DiT + Streamed Text Encoder**:
  * DiT (`best_quality_fp4.safetensors`): **4.08 GB resident VRAM**.
  * VAE (`AutoencoderKLQwenImage21`): **0.64 GB resident VRAM**.
  * Qwen3-VL-8B ViT Vision Encoder: **1.07 GB resident VRAM**.
  * Qwen3-VL-8B Language Model: Streamed layer-by-layer through a static **368 MB GPU buffer** over PCIe at ~28.7 GB/s via `stream_encoder.py`.
* **Total VRAM Footprint**: **~6.17 GB active VRAM**, leaving **~9.5 GB free headroom** on a single 16 GB GPU (such as RTX 5060 Ti or RTX 4080)!
* **Inference Speed**: Multimodal prompt encoding completes in **1.06s** (saving 16s vs CPU), and 25-step image generation runs in **~20s**.

---

## πŸš€ Quickstart Usage

### 1. Installation

```bash
pip install diffusers transformers accelerate torch sentencepiece
```

### 2. Loading with Diffusers

```python
import torch
from diffusers import QwenImage21Pipeline
from transformers import Qwen3VLForConditionalGeneration, Qwen3VLProcessor

# 1. Load Heretic text encoder and processor
text_encoder = Qwen3VLForConditionalGeneration.from_pretrained(
    "models/Qwen21_Text_Encoder_Heretic",
    torch_dtype=torch.bfloat16,
    low_cpu_mem_usage=True,
)
processor = Qwen3VLProcessor.from_pretrained("models/Qwen21_Text_Encoder_Heretic")

# 2. Assemble into pipeline
pipe = QwenImage21Pipeline.from_pretrained(
    "Qwen/Qwen-Image-2.1",
    text_encoder=text_encoder,
    processor=processor,
    torch_dtype=torch.bfloat16,
)
pipe.enable_sequential_cpu_offload(gpu_id=0)

# 3. Generate with precise prompt adherence
image = pipe(
    prompt="Two cute anime girls in colorful bikinis playing beach volleyball on a sunny tropical beach, dynamic action pose, jumping to spike the ball, sharp focus",
    height=1024,
    width=1024,
    num_inference_steps=25,
    true_cfg_scale=1.0,
).images[0]

image.save("beach_volleyball.png")
```

### 3. High-Throughput Server Usage

Run the bundled ImageEditServer with NVFP4 DiT and Heretic text encoder:

```bash
# Start server on port 4500 (uses Heretic text encoder by default)
./extras/imagegen_qwen21_nvfp4.sh 4500
```

---

## πŸ“¦ Packaged Sources (`extras/`)

* `create_heretic_text_encoder.py`: Complete script used to measure refusal vectors and perform norm-preserving biprojected abliteration.
* `stream_encoder.py`: Zero-quality-loss layerwise weight streaming engine for Qwen3-VL-8B.
* `test_heretic_beach_volleyball.py`: Empirical verification script comparing stock vs Heretic encoders.
* `QwenImage21NVFP4Backend.py`: Diffusers + Nunchaku backend supporting custom text encoder overrides.
* `ImageEditServer.py` & `imagegen_qwen21_nvfp4.sh`: Resident image generation server.

---

## πŸ“œ Citation & Credits

* **Qwen-Image-2.1 & Qwen3-VL**: Qwen Team, Alibaba Cloud.
* **Abliteration Principles**: Arditi et al. (*Refusal in Language Models Is Mediated by a Single Direction*).
* **Heretic LLM**: Heretic project (*Directional Abliteration Toolkit*).
* **Abliteration & Diffusion Conditioning Optimization**: Oleg K. / Nikola Seeker Project.