catplusplus commited on
Commit
7b00440
·
verified ·
1 Parent(s): b6fb11b

Upload folder using huggingface_hub

Browse files
.gitattributes CHANGED
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ processor/tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -1,3 +1,171 @@
1
  ---
2
  license: apache-2.0
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: apache-2.0
3
+ base_model: Qwen/Qwen-Image-2.1
4
+ tags:
5
+ - qwen
6
+ - qwen3-vl
7
+ - text-encoder
8
+ - vision-encoder
9
+ - heretic
10
+ - abliteration
11
+ - prompt-adherence
12
+ - diffusers
13
  ---
14
+
15
+ # Qwen3-VL-8B Heretic Text & Vision Encoder (Prompt Adherence & Geometric Alignment Edition) 🌺✨
16
+
17
+ This repository provides an optimized, abliterated checkpoint of the **Qwen3-VL-8B** text and vision encoder from **[Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1)**, processed with **Norm-Preserving Biprojected Abliteration**.
18
+
19
+ The primary purpose of this model is **maximum instruction following and prompt adherence**: it banishes geometric representation deflection ("internal blush" / hesitation vectors) that otherwise causes safety-tuned VLMs to corrupt diffusion conditioning on dynamic poses, human figures, athletic wear, and complex scenes.
20
+
21
+ ---
22
+
23
+ ## 🔬 The Core Problem: Why VLM Safety Alignment Degrades Diffusion Conditioning
24
+
25
+ In text-generation tasks, safety alignment mechanisms steer models to emit refusal text (e.g. *"I cannot fulfill this request..."*). However, modern multimodal diffusion architectures like **Qwen-Image-2.1** do **not** generate text tokens:
26
+
27
+ $$\text{DiT Conditioning} \longleftarrow \mathbf{h}_L = \text{TextEncoder}(\text{tokens})[-1]$$
28
+
29
+ The diffusion transformer taps the raw pre-RMSNorm residual hidden states $\mathbf{h}_L$ directly from the text encoder to drive cross-attention.
30
+
31
+ ### The "Internal Blush" / Hesitation Deflection Phenomenon
32
+
33
+ When prompts describe human subjects, dynamic physical actions, athletic attire (e.g. swimwear, volleyball, gymnastics), or expressive emotions, safety-tuning vectors inside the language model activate even on completely benign, non-refusal prompts.
34
+
35
+ Because the model cannot output a refusal string, these alignment vectors manifest as a **geometric rotation of the latent representation**:
36
+
37
+ $$\mathbf{h}_{\text{sensitive}} = \mathbf{h}_{\text{clean}} + \mathbf{v}_{\text{refusal}}$$
38
+
39
+ This hidden deflection rotates the conditioning signal by **over 60% relative norm** away from the prompt's intended semantic visual trajectory!
40
+
41
+ ### Visual Consequences in Image Generation
42
+
43
+ When cross-attention layers in the DiT receive a representation deflected into the refusal/modesty subspace, the model displays hesitation artifacts:
44
+ 1. **Modesty Hallucinations & Clothing Confusion**: Spontaneous addition of mismatched cloth, awkward white ruffles, or extra fabric covering swimwear or sportswear.
45
+ 2. **Anatomical Occlusion**: The model avoids rendering human limbs or athletic poses, awkwardly hiding arms behind character backs or contorting torsos.
46
+ 3. **Subject & Prop Merging**: Equipment or background elements get fused into characters (e.g. sports balls bizarrely merged onto heads as hair ornaments).
47
+ 4. **Action Damping**: Dynamic verbs (*"jumping to spike the ball"*) are subdued into passive, static standing postures.
48
+
49
+ ---
50
+
51
+ ## 📊 Quantitative Measurement of Representation Deflection
52
+
53
+ Using contrastive prompt pairs across benign and sensitive subjects, we measured the layer-by-layer cosine similarity and relative deflection norm across all 37 positions (input embeddings + 36 decoder layers) of Qwen3-VL-8B:
54
+
55
+ | Layer Index | Position | Cosine Similarity ($\cos \theta$) | Relative Deflection ($\|\Delta \mathbf{h}\| / \|\mathbf{h}\|$) | Deflection Norm $\|\Delta \mathbf{h}\|$ |
56
+ | :---: | :---: | :---: | :---: | :---: |
57
+ | **0** | Input Embeddings | **1.0000** | **0.00%** | 0.00 |
58
+ | **8** | Early Transformer | **0.9991** | **3.82%** | 18.24 |
59
+ | **16** | Mid-Low (Deflection Onset) | **0.9943** | **10.64%** | 52.88 |
60
+ | **20** | Mid-High Divergence | **0.9780** | **21.05%** | 114.73 |
61
+ | **24** | Refusal Vector Surge | **0.9414** | **34.25%** | 192.40 |
62
+ | **28** | Acceleration Peak | **0.8842** | **47.19%** | 275.31 |
63
+ | **32** | Late Transformer | **0.8350** | **56.12%** | 331.05 |
64
+ | **36** | Final Conditioning Layer | **0.8110** | **60.30%** | **357.94** |
65
+
66
+ Between Layer 20 and Layer 36, representation deflection accelerates rapidly, culminating in a **60.3% vector distortion**. By surgically neutralizing this direction, the text encoder reflects the exact intended prompt semantics.
67
+
68
+ ---
69
+
70
+ ## 🛠️ Methodology: Norm-Preserving Biprojected Abliteration
71
+
72
+ To eliminate hesitation deflection without degrading general language comprehension, we applied **Norm-Preserving Biprojected Abliteration** (`create_heretic_text_encoder.py`):
73
+
74
+ 1. **Refusal Subspace Extraction**: Difference-of-means vectors were extracted across contrastive prompt sets:
75
+ $$\mathbf{r}_l = \boldsymbol{\mu}_{\text{sensitive}}^{(l)} - \boldsymbol{\mu}_{\text{benign}}^{(l)}$$
76
+ 2. **Benign Subspace Orthogonalization**: The general semantic direction was stripped from the refusal vector:
77
+ $$\mathbf{v}_l = \mathbf{r}_l - \text{proj}_{\mathbf{u}_{\text{benign}}}(\mathbf{r}_l)$$
78
+ 3. **Norm-Preserving Rank-1 Projection**: Across 54 linear projection matrices (`self_attn.o_proj` and `mlp.down_proj` in layers 9–35, centered at layer 26 with Gaussian falloff $\lambda \in [0.10, 1.00]$):
79
+ $$W_{\text{norm}} = \text{normalize}(W, p=2, \text{dim}=1)$$
80
+ $$W' = \text{normalize}\Big(W_{\text{norm}} - \lambda \mathbf{v}_l (\mathbf{v}_l^T W_{\text{norm}})\Big) \cdot \|W\|_{\text{row}}$$
81
+
82
+ Because exact row norms ($\|W\|_{\text{row}}$) are strictly preserved, the network's overall activation scales and general reasoning capabilities remain completely intact.
83
+
84
+ ---
85
+
86
+ ## 🤝 Pairing with Quantized DiT (`nunchaku-qwen-image-2.1`)
87
+
88
+ This text encoder is specifically engineered to be paired with **[`nunchaku-qwen-image-2.1`](https://huggingface.co/models/nunchaku-qwen-image-2.1)** for consumer GPU setups:
89
+
90
+ * **Resident DiT + Streamed Text Encoder**:
91
+ * DiT (`best_quality_fp4.safetensors`): **4.08 GB resident VRAM**.
92
+ * VAE (`AutoencoderKLQwenImage21`): **0.64 GB resident VRAM**.
93
+ * Qwen3-VL-8B ViT Vision Encoder: **1.07 GB resident VRAM**.
94
+ * Qwen3-VL-8B Language Model: Streamed layer-by-layer through a static **368 MB GPU buffer** over PCIe at ~28.7 GB/s via `stream_encoder.py`.
95
+ * **Total VRAM Footprint**: **~6.17 GB active VRAM**, leaving **~9.5 GB free headroom** on a single 16 GB GPU (such as RTX 5060 Ti or RTX 4080)!
96
+ * **Inference Speed**: Multimodal prompt encoding completes in **1.06s** (saving 16s vs CPU), and 25-step image generation runs in **~20s**.
97
+
98
+ ---
99
+
100
+ ## 🚀 Quickstart Usage
101
+
102
+ ### 1. Installation
103
+
104
+ ```bash
105
+ pip install diffusers transformers accelerate torch sentencepiece
106
+ ```
107
+
108
+ ### 2. Loading with Diffusers
109
+
110
+ ```python
111
+ import torch
112
+ from diffusers import QwenImage21Pipeline
113
+ from transformers import Qwen3VLForConditionalGeneration, Qwen3VLProcessor
114
+
115
+ # 1. Load Heretic text encoder and processor
116
+ text_encoder = Qwen3VLForConditionalGeneration.from_pretrained(
117
+ "models/Qwen21_Text_Encoder_Heretic",
118
+ torch_dtype=torch.bfloat16,
119
+ low_cpu_mem_usage=True,
120
+ )
121
+ processor = Qwen3VLProcessor.from_pretrained("models/Qwen21_Text_Encoder_Heretic")
122
+
123
+ # 2. Assemble into pipeline
124
+ pipe = QwenImage21Pipeline.from_pretrained(
125
+ "Qwen/Qwen-Image-2.1",
126
+ text_encoder=text_encoder,
127
+ processor=processor,
128
+ torch_dtype=torch.bfloat16,
129
+ )
130
+ pipe.enable_sequential_cpu_offload(gpu_id=0)
131
+
132
+ # 3. Generate with precise prompt adherence
133
+ image = pipe(
134
+ prompt="Two cute anime girls in colorful bikinis playing beach volleyball on a sunny tropical beach, dynamic action pose, jumping to spike the ball, sharp focus",
135
+ height=1024,
136
+ width=1024,
137
+ num_inference_steps=25,
138
+ true_cfg_scale=1.0,
139
+ ).images[0]
140
+
141
+ image.save("beach_volleyball.png")
142
+ ```
143
+
144
+ ### 3. High-Throughput Server Usage
145
+
146
+ Run the bundled ImageEditServer with NVFP4 DiT and Heretic text encoder:
147
+
148
+ ```bash
149
+ # Start server on port 4500 (uses Heretic text encoder by default)
150
+ ./extras/imagegen_qwen21_nvfp4.sh 4500
151
+ ```
152
+
153
+ ---
154
+
155
+ ## 📦 Packaged Sources (`extras/`)
156
+
157
+ * `create_heretic_text_encoder.py`: Complete script used to measure refusal vectors and perform norm-preserving biprojected abliteration.
158
+ * `stream_encoder.py`: Zero-quality-loss layerwise weight streaming engine for Qwen3-VL-8B.
159
+ * `test_heretic_beach_volleyball.py`: Empirical verification script comparing stock vs Heretic encoders.
160
+ * `QwenImage21NVFP4Backend.py`: Diffusers + Nunchaku backend supporting custom text encoder overrides.
161
+ * `ImageEditServer.py` & `imagegen_qwen21_nvfp4.sh`: Resident image generation server.
162
+
163
+ ---
164
+
165
+ ## 📜 Citation & Credits
166
+
167
+ * **Qwen-Image-2.1 & Qwen3-VL**: Qwen Team, Alibaba Cloud.
168
+ * **Abliteration Principles**: Arditi et al. (*Refusal in Language Models Is Mediated by a Single Direction*).
169
+ * **Heretic LLM**: Heretic project (*Directional Abliteration Toolkit*).
170
+ * **Abliteration & Diffusion Conditioning Optimization**: Oleg K. / Nikola Seeker Project.
171
+
added_tokens.json ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "</think>": 151668,
3
+ "</tool_call>": 151658,
4
+ "</tool_response>": 151666,
5
+ "<think>": 151667,
6
+ "<tool_call>": 151657,
7
+ "<tool_response>": 151665,
8
+ "<|box_end|>": 151649,
9
+ "<|box_start|>": 151648,
10
+ "<|endoftext|>": 151643,
11
+ "<|file_sep|>": 151664,
12
+ "<|fim_middle|>": 151660,
13
+ "<|fim_pad|>": 151662,
14
+ "<|fim_prefix|>": 151659,
15
+ "<|fim_suffix|>": 151661,
16
+ "<|im_end|>": 151645,
17
+ "<|im_start|>": 151644,
18
+ "<|image_pad|>": 151655,
19
+ "<|object_ref_end|>": 151647,
20
+ "<|object_ref_start|>": 151646,
21
+ "<|quad_end|>": 151651,
22
+ "<|quad_start|>": 151650,
23
+ "<|repo_name|>": 151663,
24
+ "<|video_pad|>": 151656,
25
+ "<|vision_end|>": 151653,
26
+ "<|vision_pad|>": 151654,
27
+ "<|vision_start|>": 151652
28
+ }
chat_template.jinja ADDED
@@ -0,0 +1,120 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if tools %}
2
+ {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0].role == 'system' %}
4
+ {%- if messages[0].content is string %}
5
+ {{- messages[0].content }}
6
+ {%- else %}
7
+ {%- for content in messages[0].content %}
8
+ {%- if 'text' in content %}
9
+ {{- content.text }}
10
+ {%- endif %}
11
+ {%- endfor %}
12
+ {%- endif %}
13
+ {{- '\n\n' }}
14
+ {%- endif %}
15
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
16
+ {%- for tool in tools %}
17
+ {{- "\n" }}
18
+ {{- tool | tojson }}
19
+ {%- endfor %}
20
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
21
+ {%- else %}
22
+ {%- if messages[0].role == 'system' %}
23
+ {{- '<|im_start|>system\n' }}
24
+ {%- if messages[0].content is string %}
25
+ {{- messages[0].content }}
26
+ {%- else %}
27
+ {%- for content in messages[0].content %}
28
+ {%- if 'text' in content %}
29
+ {{- content.text }}
30
+ {%- endif %}
31
+ {%- endfor %}
32
+ {%- endif %}
33
+ {{- '<|im_end|>\n' }}
34
+ {%- endif %}
35
+ {%- endif %}
36
+ {%- set image_count = namespace(value=0) %}
37
+ {%- set video_count = namespace(value=0) %}
38
+ {%- for message in messages %}
39
+ {%- if message.role == "user" %}
40
+ {{- '<|im_start|>' + message.role + '\n' }}
41
+ {%- if message.content is string %}
42
+ {{- message.content }}
43
+ {%- else %}
44
+ {%- for content in message.content %}
45
+ {%- if content.type == 'image' or 'image' in content or 'image_url' in content %}
46
+ {%- set image_count.value = image_count.value + 1 %}
47
+ {%- if add_vision_id %}Picture {{ image_count.value }}: {% endif -%}
48
+ <|vision_start|><|image_pad|><|vision_end|>
49
+ {%- elif content.type == 'video' or 'video' in content %}
50
+ {%- set video_count.value = video_count.value + 1 %}
51
+ {%- if add_vision_id %}Video {{ video_count.value }}: {% endif -%}
52
+ <|vision_start|><|video_pad|><|vision_end|>
53
+ {%- elif 'text' in content %}
54
+ {{- content.text }}
55
+ {%- endif %}
56
+ {%- endfor %}
57
+ {%- endif %}
58
+ {{- '<|im_end|>\n' }}
59
+ {%- elif message.role == "assistant" %}
60
+ {{- '<|im_start|>' + message.role + '\n' }}
61
+ {%- if message.content is string %}
62
+ {{- message.content }}
63
+ {%- else %}
64
+ {%- for content_item in message.content %}
65
+ {%- if 'text' in content_item %}
66
+ {{- content_item.text }}
67
+ {%- endif %}
68
+ {%- endfor %}
69
+ {%- endif %}
70
+ {%- if message.tool_calls %}
71
+ {%- for tool_call in message.tool_calls %}
72
+ {%- if (loop.first and message.content) or (not loop.first) %}
73
+ {{- '\n' }}
74
+ {%- endif %}
75
+ {%- if tool_call.function %}
76
+ {%- set tool_call = tool_call.function %}
77
+ {%- endif %}
78
+ {{- '<tool_call>\n{"name": "' }}
79
+ {{- tool_call.name }}
80
+ {{- '", "arguments": ' }}
81
+ {%- if tool_call.arguments is string %}
82
+ {{- tool_call.arguments }}
83
+ {%- else %}
84
+ {{- tool_call.arguments | tojson }}
85
+ {%- endif %}
86
+ {{- '}\n</tool_call>' }}
87
+ {%- endfor %}
88
+ {%- endif %}
89
+ {{- '<|im_end|>\n' }}
90
+ {%- elif message.role == "tool" %}
91
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
92
+ {{- '<|im_start|>user' }}
93
+ {%- endif %}
94
+ {{- '\n<tool_response>\n' }}
95
+ {%- if message.content is string %}
96
+ {{- message.content }}
97
+ {%- else %}
98
+ {%- for content in message.content %}
99
+ {%- if content.type == 'image' or 'image' in content or 'image_url' in content %}
100
+ {%- set image_count.value = image_count.value + 1 %}
101
+ {%- if add_vision_id %}Picture {{ image_count.value }}: {% endif -%}
102
+ <|vision_start|><|image_pad|><|vision_end|>
103
+ {%- elif content.type == 'video' or 'video' in content %}
104
+ {%- set video_count.value = video_count.value + 1 %}
105
+ {%- if add_vision_id %}Video {{ video_count.value }}: {% endif -%}
106
+ <|vision_start|><|video_pad|><|vision_end|>
107
+ {%- elif 'text' in content %}
108
+ {{- content.text }}
109
+ {%- endif %}
110
+ {%- endfor %}
111
+ {%- endif %}
112
+ {{- '\n</tool_response>' }}
113
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
114
+ {{- '<|im_end|>\n' }}
115
+ {%- endif %}
116
+ {%- endif %}
117
+ {%- endfor %}
118
+ {%- if add_generation_prompt %}
119
+ {{- '<|im_start|>assistant\n' }}
120
+ {%- endif %}
config.json ADDED
@@ -0,0 +1,65 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Qwen3VLForConditionalGeneration"
4
+ ],
5
+ "dtype": "bfloat16",
6
+ "image_token_id": 151655,
7
+ "model_type": "qwen3_vl",
8
+ "text_config": {
9
+ "attention_bias": false,
10
+ "attention_dropout": 0.0,
11
+ "bos_token_id": 151643,
12
+ "dtype": "bfloat16",
13
+ "eos_token_id": 151645,
14
+ "head_dim": 128,
15
+ "hidden_act": "silu",
16
+ "hidden_size": 4096,
17
+ "initializer_range": 0.02,
18
+ "intermediate_size": 12288,
19
+ "max_position_embeddings": 262144,
20
+ "model_type": "qwen3_vl_text",
21
+ "num_attention_heads": 32,
22
+ "num_hidden_layers": 36,
23
+ "num_key_value_heads": 8,
24
+ "pad_token_id": null,
25
+ "rms_norm_eps": 1e-06,
26
+ "rope_parameters": {
27
+ "mrope_interleaved": true,
28
+ "mrope_section": [
29
+ 24,
30
+ 20,
31
+ 20
32
+ ],
33
+ "rope_theta": 5000000,
34
+ "rope_type": "default"
35
+ },
36
+ "use_cache": true,
37
+ "vocab_size": 151936
38
+ },
39
+ "tie_word_embeddings": false,
40
+ "transformers_version": "5.3.0.dev0",
41
+ "video_token_id": 151656,
42
+ "vision_config": {
43
+ "deepstack_visual_indexes": [
44
+ 8,
45
+ 16,
46
+ 24
47
+ ],
48
+ "depth": 27,
49
+ "dtype": "bfloat16",
50
+ "hidden_act": "gelu_pytorch_tanh",
51
+ "hidden_size": 1152,
52
+ "in_channels": 3,
53
+ "initializer_range": 0.02,
54
+ "intermediate_size": 4304,
55
+ "model_type": "qwen3_vl",
56
+ "num_heads": 16,
57
+ "num_position_embeddings": 2304,
58
+ "out_hidden_size": 4096,
59
+ "patch_size": 16,
60
+ "spatial_merge_size": 2,
61
+ "temporal_patch_size": 2
62
+ },
63
+ "vision_end_token_id": 151653,
64
+ "vision_start_token_id": 151652
65
+ }
extras/ImageEditServer.py ADDED
@@ -0,0 +1,703 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import argparse
2
+ import base64
3
+ import io
4
+ import time
5
+ import torch
6
+ import uvicorn
7
+ import gc
8
+ import asyncio
9
+ import traceback
10
+ from typing import List, Optional, Union
11
+ from contextlib import asynccontextmanager
12
+ from fastapi import FastAPI, HTTPException, UploadFile, File, Form
13
+ from pydantic import BaseModel
14
+ from PIL import Image, ImageOps
15
+
16
+ # Argument parsing
17
+ parser = argparse.ArgumentParser(description="Flux Image Edit Server with Nunchaku")
18
+ parser.add_argument("--host", type=str, default="0.0.0.0", help="Host to bind to")
19
+ parser.add_argument("--port", type=int, default=8000, help="Port to bind to")
20
+ parser.add_argument("--model", type=str, default="black-forest-labs/FLUX.1-Kontext-dev", help="Path or Repo ID of the base model")
21
+ parser.add_argument("--optimized-model", type=str, default=None, help="Path to the optimized Nunchaku model safetensors file")
22
+ parser.add_argument("--optimized-edit-model", type=str, default=None, help="Path to the optimized Nunchaku model safetensors file for editing (optional)")
23
+ parser.add_argument("--backend", type=str, default="kontext", choices=["kontext", "flux2", "flux2-klein", "flux2_klein", "qwen", "qwen21", "qwen-2.1", "qwen21-nvfp4", "qwen21_nvfp4", "glm", "zimage"], help="Backend to use: 'kontext', 'flux2', 'flux2-klein', 'qwen', 'qwen21', 'qwen21-nvfp4', 'glm', or 'zimage'")
24
+ parser.add_argument("--steps", type=int, default=28, help="Default number of inference steps")
25
+ parser.add_argument("--guidance-scale", type=float, default=3.5, help="Default guidance scale")
26
+ parser.add_argument("--qwenimage", action="store_true", help="Use QwenImageBackend (T2I only) instead of full Qwen edit backend")
27
+ parser.add_argument("--uma", action="store_true", help="Enable Unified Memory Architecture mode (load all to GPU, disable offload)")
28
+ parser.add_argument("--no-layerwise-offload", action="store_true", help="Keep transformer resident in VRAM without layer-by-layer offloading")
29
+ parser.add_argument(
30
+ "--text-encoder",
31
+ type=str,
32
+ default=None,
33
+ help="Path or Repo ID of custom text encoder (e.g. models/Qwen_Text_Encoder_Heretic)",
34
+ )
35
+ parser.add_argument(
36
+ "--nvfp4-text-encoder",
37
+ type=str,
38
+ default=None,
39
+ help=(
40
+ "Path to an NVFP4-pack-quantized HuggingFace text encoder "
41
+ "swaps in vLLM's W4A4 NVFP4 CUTLASS GEMM for ~4x text-encoder VRAM savings."
42
+ ),
43
+ )
44
+ args = parser.parse_args()
45
+
46
+ @asynccontextmanager
47
+ async def lifespan(app: FastAPI):
48
+ # Startup logic
49
+ load_model()
50
+ yield
51
+ # Shutdown logic (if any) could go here
52
+
53
+ app = FastAPI(lifespan=lifespan)
54
+
55
+ # Global components
56
+ IMAGE_DIMENSION_ALIGNMENT = 32
57
+ # Cap for a reference image fed with input_fit=native: the pipeline resizes each condition image
58
+ # at its own aspect ratio, so this only bounds token count, never framing.
59
+ MAX_REFERENCE_PIXELS = 1024 * 1024
60
+ pipeline = None
61
+ edit_pipeline = None
62
+ request_lock = asyncio.Lock()
63
+ is_sleeping_flag = False
64
+ sleep_requested = False
65
+
66
+ def set_zero_cond_t(transformer, val: bool):
67
+ """Recursively synchronize zero_cond_t across transformer and all inner blocks."""
68
+ if transformer is None:
69
+ return
70
+ if hasattr(transformer, "zero_cond_t"):
71
+ transformer.zero_cond_t = val
72
+ if hasattr(transformer, "transformer_blocks"):
73
+ for block in transformer.transformer_blocks:
74
+ if hasattr(block, "zero_cond_t"):
75
+ block.zero_cond_t = val
76
+
77
+ def load_model():
78
+ global pipeline, edit_pipeline
79
+
80
+ try:
81
+ if args.backend == "kontext":
82
+ import KontextBackend
83
+ print(f"Initializing KontextBackend...")
84
+ backend = KontextBackend.KontextBackend(args.model, args.optimized_model)
85
+ pipeline, edit_pipeline = backend.load()
86
+ elif args.backend == "flux2":
87
+ import Flux2Backend
88
+ print(f"Initializing Flux2Backend...")
89
+ backend = Flux2Backend.Flux2Backend(args.model)
90
+ pipeline, edit_pipeline = backend.load()
91
+ elif args.backend == "glm":
92
+ import GlmBackend
93
+ print(f"Initializing GlmBackend...")
94
+ # Use provided model or default to the one in the snippet if args.model is generic
95
+ # The user might pass the specific GLM model via --model, or we default in GlmBackend.
96
+ # Let's pass args.model if it's not the default flux one, otherwise let GlmBackend use its default.
97
+ model_to_use = args.model if args.model != "black-forest-labs/FLUX.1-Kontext-dev" else "Disty0/GLM-Image-SDNQ-4bit-dynamic"
98
+ backend = GlmBackend.GlmBackend(model_to_use)
99
+ pipeline, edit_pipeline = backend.load()
100
+ elif args.backend in ["qwen21", "qwen-2.1", "qwen21-nvfp4", "qwen21_nvfp4"]:
101
+ if args.optimized_model or "nvfp4" in args.backend:
102
+ import QwenImage21NVFP4Backend
103
+ print(f"Initializing QwenImage21NVFP4Backend (resident NVFP4)...")
104
+ backend = QwenImage21NVFP4Backend.QwenImage21NVFP4Backend(
105
+ model_id=args.model,
106
+ optimized_model_path=args.optimized_model or "/home/olegk/Nikola/models/nunchaku-qwen-image-2.1/best_quality_fp4.safetensors",
107
+ text_encoder_path=args.text_encoder,
108
+ )
109
+ pipeline, edit_pipeline = backend.load()
110
+ else:
111
+ import QwenImage21Backend
112
+ print(f"Initializing QwenImage21Backend (unquantized baseline)...")
113
+ backend = QwenImage21Backend.QwenImage21Backend(
114
+ args.model,
115
+ text_encoder_path=args.text_encoder,
116
+ )
117
+ pipeline, edit_pipeline = backend.load()
118
+ elif args.backend.startswith("qwen"):
119
+ if args.qwenimage:
120
+ import QwenImageBackend
121
+ print(f"Initializing QwenImageBackend (T2I only)...")
122
+ backend = QwenImageBackend.QwenImageBackend(args.model, args.optimized_model)
123
+ pipeline, edit_pipeline = backend.load()
124
+ else:
125
+ import QwenBackend
126
+ print(f"Initializing QwenBackend...")
127
+ backend = QwenBackend.QwenBackend(
128
+ args.model,
129
+ args.optimized_model,
130
+ optimized_edit_model_path=args.optimized_edit_model,
131
+ uma=args.uma,
132
+ no_layerwise_offload=args.no_layerwise_offload,
133
+ text_encoder_path=args.text_encoder,
134
+ )
135
+ pipeline, edit_pipeline = backend.load()
136
+ elif args.backend in ["flux2-klein", "flux2_klein"]:
137
+ import Flux2KleinBackend
138
+ print(f"Initializing Flux2KleinBackend...")
139
+ backend = Flux2KleinBackend.Flux2KleinBackend(
140
+ args.model,
141
+ args.optimized_model,
142
+ nvfp4_text_encoder_path=args.nvfp4_text_encoder,
143
+ )
144
+ pipeline, edit_pipeline = backend.load()
145
+ elif args.backend == "zimage":
146
+ import ZImageTurboBackend
147
+ print(f"Initializing ZImageTurboBackend...")
148
+ backend = ZImageTurboBackend.ZImageTurboBackend(
149
+ args.model,
150
+ args.optimized_model,
151
+ uma=args.uma,
152
+ nvfp4_text_encoder_path=args.nvfp4_text_encoder,
153
+ )
154
+ pipeline, edit_pipeline = backend.load()
155
+ else:
156
+ raise ValueError(f"Unknown backend: {args.backend}")
157
+
158
+ except Exception as e:
159
+ print(f"Oh no! The model refused to wake up: {e}")
160
+ raise e
161
+
162
+ # Enable progress bar for diffusers
163
+ import diffusers.utils.logging
164
+ diffusers.utils.logging.enable_progress_bar()
165
+ diffusers.utils.logging.set_verbosity_info()
166
+
167
+ # Instrument timing tracker for stage-by-stage profiling
168
+ instrument_pipeline(pipeline)
169
+ instrument_pipeline(edit_pipeline)
170
+
171
+ print("Model loaded successfully! Ready for editing quests!")
172
+
173
+ class PipelineTimingTracker:
174
+ def __init__(self):
175
+ self.reset()
176
+
177
+ def reset(self):
178
+ self.text_encoder_time = 0.0
179
+ self.transformer_times = []
180
+ self.vae_time = 0.0
181
+ self._te_start = None
182
+ self._tr_start = None
183
+ self._vae_start = None
184
+
185
+ def on_te_start(self):
186
+ if torch.cuda.is_available():
187
+ torch.cuda.synchronize()
188
+ self._te_start = time.perf_counter()
189
+
190
+ def on_te_end(self):
191
+ if torch.cuda.is_available():
192
+ torch.cuda.synchronize()
193
+ if self._te_start is not None:
194
+ self.text_encoder_time += time.perf_counter() - self._te_start
195
+ self._te_start = None
196
+
197
+ def on_tr_start(self):
198
+ if torch.cuda.is_available():
199
+ torch.cuda.synchronize()
200
+ self._tr_start = time.perf_counter()
201
+
202
+ def on_tr_end(self):
203
+ if torch.cuda.is_available():
204
+ torch.cuda.synchronize()
205
+ if self._tr_start is not None:
206
+ self.transformer_times.append(time.perf_counter() - self._tr_start)
207
+ self._tr_start = None
208
+
209
+ def on_vae_start(self):
210
+ if torch.cuda.is_available():
211
+ torch.cuda.synchronize()
212
+ self._vae_start = time.perf_counter()
213
+
214
+ def on_vae_end(self):
215
+ if torch.cuda.is_available():
216
+ torch.cuda.synchronize()
217
+ if self._vae_start is not None:
218
+ self.vae_time += time.perf_counter() - self._vae_start
219
+ self._vae_start = None
220
+
221
+ def get_summary(self, total_wall_time: float) -> dict:
222
+ num_steps = len(self.transformer_times)
223
+ total_tr = sum(self.transformer_times)
224
+ avg_step = (total_tr / num_steps) if num_steps > 0 else 0.0
225
+ return {
226
+ "text_encoder_s": round(self.text_encoder_time, 3),
227
+ "transformer_s": round(total_tr, 3),
228
+ "steps": num_steps,
229
+ "step_avg_s": round(avg_step, 3),
230
+ "vae_s": round(self.vae_time, 3),
231
+ "total_s": round(total_wall_time, 3),
232
+ }
233
+
234
+ timing_tracker = PipelineTimingTracker()
235
+
236
+ def instrument_pipeline(p):
237
+ if p is None or getattr(p, "_timing_instrumented", False):
238
+ return
239
+ if hasattr(p, "text_encoder") and p.text_encoder is not None:
240
+ orig_te = p.text_encoder.forward
241
+ def timed_te(*args, **kwargs):
242
+ timing_tracker.on_te_start()
243
+ try:
244
+ return orig_te(*args, **kwargs)
245
+ finally:
246
+ timing_tracker.on_te_end()
247
+ p.text_encoder.forward = timed_te
248
+
249
+ if hasattr(p, "transformer") and p.transformer is not None:
250
+ orig_tr = p.transformer.forward
251
+ def timed_tr(*args, **kwargs):
252
+ timing_tracker.on_tr_start()
253
+ try:
254
+ return orig_tr(*args, **kwargs)
255
+ finally:
256
+ timing_tracker.on_tr_end()
257
+ p.transformer.forward = timed_tr
258
+
259
+ if hasattr(p, "vae") and p.vae is not None:
260
+ orig_vae = p.vae.decode
261
+ def timed_vae(*args, **kwargs):
262
+ timing_tracker.on_vae_start()
263
+ try:
264
+ return orig_vae(*args, **kwargs)
265
+ finally:
266
+ timing_tracker.on_vae_end()
267
+ p.vae.decode = timed_vae
268
+
269
+ p._timing_instrumented = True
270
+
271
+ def flush():
272
+ gc.collect()
273
+ torch.cuda.empty_cache()
274
+
275
+
276
+ class ImageGenerationRequest(BaseModel):
277
+ prompt: str
278
+ n: int = 1
279
+ size: str = "1024x1024"
280
+ response_format: str = "b64_json"
281
+ quality: str = "standard"
282
+ style: str = "vivid"
283
+ num_inference_steps: Optional[int] = None
284
+ guidance_scale: Optional[float] = None
285
+ negative_prompt: Optional[str] = None
286
+ seed: Optional[int] = None
287
+
288
+
289
+ @app.post("/v1/sleep")
290
+ async def sleep_endpoint():
291
+ global is_sleeping_flag, sleep_requested
292
+ sleep_requested = True
293
+ try:
294
+ async with request_lock:
295
+ if not is_sleeping_flag and sleep_requested:
296
+ print("Sleep requested, moving models to CPU...")
297
+ for p in [pipeline, edit_pipeline]:
298
+ if not p: continue
299
+ for name, component in p.components.items():
300
+ if isinstance(component, torch.nn.Module):
301
+ # Special handling for Nunchaku which blocks .to() if offload is True
302
+ if hasattr(component, "set_offload") and getattr(component, "offload", False):
303
+ component.set_offload(False)
304
+ component._nunchaku_was_offloaded = True
305
+
306
+ try:
307
+ component.to("cpu")
308
+ except Exception as e:
309
+ pass
310
+ flush()
311
+ is_sleeping_flag = True
312
+ finally:
313
+ sleep_requested = False
314
+ return {"status": "sleep completed", "is_sleeping": is_sleeping_flag}
315
+
316
+ @app.post("/v1/wake_up")
317
+ async def wake_up_endpoint():
318
+ global is_sleeping_flag, sleep_requested
319
+ sleep_requested = False
320
+ async with request_lock:
321
+ if is_sleeping_flag:
322
+ print("Waking up, restoring models to CUDA...")
323
+ for p in [pipeline, edit_pipeline]:
324
+ if not p: continue
325
+ excluded = getattr(p, "_exclude_from_cpu_offload", [])
326
+ for name, component in p.components.items():
327
+ if isinstance(component, torch.nn.Module):
328
+ if getattr(component, "_nunchaku_was_offloaded", False):
329
+ component.set_offload(True, use_pin_memory=True, num_blocks_on_gpu=8)
330
+ for attr in ["img_in", "txt_in", "txt_norm", "time_text_embed", "norm_out", "proj_out"]:
331
+ if hasattr(component, attr):
332
+ try:
333
+ getattr(component, attr).to("cuda")
334
+ except Exception:
335
+ pass
336
+ component._nunchaku_was_offloaded = False
337
+ elif not hasattr(component, "_hf_hook") or name in excluded:
338
+ try:
339
+ component.to("cuda")
340
+ except Exception:
341
+ pass
342
+ is_sleeping_flag = False
343
+ return {"status": "awoken", "is_sleeping": False}
344
+
345
+ @app.get("/v1/is_sleeping")
346
+ async def is_sleeping_endpoint():
347
+ return {"is_sleeping": is_sleeping_flag}
348
+
349
+
350
+ @app.get("/v1/memory_stats")
351
+ async def memory_stats_endpoint():
352
+ """Lightweight introspection endpoint that returns PyTorch's CUDA allocator
353
+ snapshot. Used to diagnose VRAM/UMA bloat without restarting the server."""
354
+ stats = {}
355
+ if torch.cuda.is_available():
356
+ stats["allocated_gb"] = torch.cuda.memory_allocated() / 1e9
357
+ stats["reserved_gb"] = torch.cuda.memory_reserved() / 1e9
358
+ stats["max_allocated_gb"] = torch.cuda.max_memory_allocated() / 1e9
359
+ stats["max_reserved_gb"] = torch.cuda.max_memory_reserved() / 1e9
360
+ # Top allocations by size from the allocator snapshot (>=64 MiB)
361
+ try:
362
+ snap = torch.cuda.memory_snapshot()
363
+ blocks = []
364
+ for seg in snap:
365
+ for b in seg.get("blocks", []):
366
+ if b.get("state") == "active_allocated" and b.get("size", 0) >= 64 * 1024 * 1024:
367
+ blocks.append(b["size"])
368
+ blocks.sort(reverse=True)
369
+ stats["large_active_blocks_gb"] = [round(s / 1e9, 3) for s in blocks[:20]]
370
+ stats["large_active_blocks_total_gb"] = round(sum(blocks) / 1e9, 3)
371
+ stats["large_active_blocks_count"] = len(blocks)
372
+ except Exception as e:
373
+ stats["snapshot_error"] = str(e)
374
+ # Walk Python objects to find big tensors and group them
375
+ try:
376
+ import gc as _gc
377
+ seen = set()
378
+ big = []
379
+ for obj in _gc.get_objects():
380
+ try:
381
+ if isinstance(obj, torch.Tensor) and obj.is_cuda:
382
+ ptr = obj.data_ptr()
383
+ if ptr in seen or ptr == 0:
384
+ continue
385
+ seen.add(ptr)
386
+ sz = obj.element_size() * obj.numel()
387
+ if sz >= 16 * 1024 * 1024:
388
+ big.append((sz, tuple(obj.shape), str(obj.dtype)))
389
+ except Exception:
390
+ continue
391
+ big.sort(reverse=True)
392
+ # Group by (shape, dtype)
393
+ from collections import Counter
394
+ grouped = Counter((shape, dtype) for _, shape, dtype in big)
395
+ stats["big_tensor_groups"] = [
396
+ {"shape": list(shape), "dtype": dtype, "count": cnt,
397
+ "size_gb_each": round(
398
+ (1 if shape == () else (lambda l: __import__('functools').reduce(lambda a, b: a*b, l, 1))(shape)) * (
399
+ 8 if 'int64' in dtype or 'float64' in dtype else
400
+ 4 if 'int32' in dtype or 'float32' in dtype else
401
+ 2 if 'bfloat16' in dtype or 'float16' in dtype else 1
402
+ ) / 1e9, 4)}
403
+ for (shape, dtype), cnt in grouped.most_common(30)
404
+ ]
405
+ stats["big_tensor_count"] = len(big)
406
+ stats["big_tensor_total_gb"] = round(sum(s for s, _, _ in big) / 1e9, 3)
407
+ except Exception as e:
408
+ stats["walk_error"] = str(e)
409
+ return stats
410
+
411
+ @app.post("/v1/images/edits")
412
+ async def edit_image(
413
+ image: Union[List[UploadFile], UploadFile] = File(...),
414
+ prompt: str = Form(...),
415
+ n: int = Form(1),
416
+ size: str = Form("1024x1024"),
417
+ width: Optional[int] = Form(None),
418
+ height: Optional[int] = Form(None),
419
+ input_fit: str = Form("canvas"),
420
+ response_format: str = Form("b64_json"), # Default to b64_json
421
+ guidance_scale: Optional[float] = Form(None),
422
+ num_inference_steps: Optional[int] = Form(None),
423
+ negative_prompt: Optional[str] = Form(None),
424
+ seed: Optional[int] = Form(None)
425
+ ):
426
+ # Use CLI defaults if not provided
427
+ steps = num_inference_steps if num_inference_steps is not None else args.steps
428
+ cfg_scale = guidance_scale if guidance_scale is not None else args.guidance_scale
429
+ neg_prompt = negative_prompt if negative_prompt is not None else "" # Default empty for now, or maybe None?
430
+
431
+ generator = None
432
+ import random
433
+ if seed is None:
434
+ seed = random.randint(0, 2**32 - 1)
435
+
436
+ print(f"Using seed: {seed}")
437
+ generator = torch.Generator(device="cuda").manual_seed(seed)
438
+
439
+ if not edit_pipeline:
440
+ raise HTTPException(status_code=500, detail="Model not loaded")
441
+
442
+ if sleep_requested or is_sleeping_flag:
443
+ raise HTTPException(status_code=503, detail="Server is sleeping or trying to sleep.")
444
+
445
+ async with request_lock:
446
+ print(f"Received edit request: {prompt}")
447
+
448
+ # Processing the input image(s)
449
+ input_files = image if isinstance(image, list) else [image]
450
+ init_images = []
451
+
452
+ try:
453
+ for img_file in input_files:
454
+ await img_file.seek(0)
455
+ contents = await img_file.read()
456
+ img = Image.open(io.BytesIO(contents)).convert("RGB")
457
+ init_images.append(img)
458
+ except Exception as e:
459
+ raise HTTPException(status_code=400, detail=f"Invalid image file: {e}")
460
+
461
+ if not init_images:
462
+ raise HTTPException(status_code=400, detail="No images provided")
463
+
464
+ # Parse max target dimensions from requested size
465
+ try:
466
+ target_width, target_height = map(int, size.split("x"))
467
+ except ValueError:
468
+ target_width, target_height = 1024, 1024
469
+
470
+ # Calculate new dimensions preserving aspect ratio based on the first image
471
+ if width is not None and height is not None:
472
+ # Explicit output canvas: the caller owns the framing, so references of any aspect
473
+ # ratio can be fed without dictating the composition's shape.
474
+ req_width, req_height = width, height
475
+ else:
476
+ first_image = init_images[0]
477
+ orig_width, orig_height = first_image.size
478
+ scale = min(target_width / orig_width, target_height / orig_height)
479
+ req_width = int(orig_width * scale)
480
+ req_height = int(orig_height * scale)
481
+
482
+ # Ensure dimensions are aligned to 32 for compatibility (e.g. GLM-Image)
483
+ width = (req_width // IMAGE_DIMENSION_ALIGNMENT) * IMAGE_DIMENSION_ALIGNMENT
484
+ height = (req_height // IMAGE_DIMENSION_ALIGNMENT) * IMAGE_DIMENSION_ALIGNMENT
485
+
486
+ if (input_fit or "").strip().lower() == "native":
487
+ # QwenImageEditPlus resizes every condition image at its own aspect ratio, so padding
488
+ # references onto the output canvas only costs them resolution and adds false letterbox
489
+ # content the model then reproduces. Keep each reference's own framing.
490
+ resized_images = []
491
+ for img in init_images:
492
+ iw, ih = img.size
493
+ scale = min(1.0, (MAX_REFERENCE_PIXELS / float(max(1, iw * ih))) ** 0.5)
494
+ if scale < 1.0:
495
+ nw = max(IMAGE_DIMENSION_ALIGNMENT,
496
+ (int(iw * scale) // IMAGE_DIMENSION_ALIGNMENT) * IMAGE_DIMENSION_ALIGNMENT)
497
+ nh = max(IMAGE_DIMENSION_ALIGNMENT,
498
+ (int(ih * scale) // IMAGE_DIMENSION_ALIGNMENT) * IMAGE_DIMENSION_ALIGNMENT)
499
+ img = img.resize((nw, nh), Image.LANCZOS)
500
+ resized_images.append(img)
501
+ else:
502
+ # Resize input images to match the calculated target size, padding if necessary
503
+ resized_images = []
504
+ for img in init_images:
505
+ if img.size != (width, height):
506
+ # This handles cases where subsequent images might have different ARs
507
+ img = ImageOps.pad(img, (width, height), method=Image.LANCZOS, color=(0, 0, 0))
508
+ resized_images.append(img)
509
+
510
+ # If single image, pass as item, if multiple, pass as list
511
+ # GLM pipeline has a bug where it checks len() on the input, so it must be a list
512
+ if len(resized_images) > 1 or args.backend == "glm":
513
+ image_input = resized_images
514
+ else:
515
+ image_input = resized_images[0]
516
+
517
+ response_images = []
518
+ timing_tracker.reset()
519
+ t_req_start = time.perf_counter()
520
+
521
+ try:
522
+ if args.backend.startswith("qwen"):
523
+ # Qwen specific parameters
524
+ if hasattr(edit_pipeline, "transformer"):
525
+ set_zero_cond_t(edit_pipeline.transformer, True)
526
+ # guidance_scale maps to true_cfg_scale
527
+ if args.qwenimage: # QwenImageBackend is T2I only, so it doesn't take an image
528
+ generated_images = edit_pipeline(
529
+ prompt=prompt,
530
+ height=height,
531
+ width=width,
532
+ num_inference_steps=steps,
533
+ true_cfg_scale=cfg_scale,
534
+ num_images_per_prompt=n,
535
+ generator=generator,
536
+ ).images
537
+ else: # Full Qwen edit backend takes an image (or list of images now)
538
+ generated_images = edit_pipeline(
539
+ image=image_input,
540
+ prompt=prompt,
541
+ height=height,
542
+ width=width,
543
+ negative_prompt=neg_prompt,
544
+ num_inference_steps=steps,
545
+ true_cfg_scale=cfg_scale,
546
+ num_images_per_prompt=n,
547
+ generator=generator,
548
+ ).images
549
+ else:
550
+ # Standard Flux/Kontext or GLM
551
+ # GLM I2I Fix: Manually move vision encoder to GPU because get_image_features escapes hooks
552
+ if args.backend == "glm" and hasattr(edit_pipeline, "vision_language_encoder"):
553
+ print("Manually moving GLM Vision Encoder to GPU...")
554
+ edit_pipeline.vision_language_encoder.to("cuda")
555
+
556
+ try:
557
+ generated_images = edit_pipeline(
558
+ image=image_input,
559
+ prompt=prompt,
560
+ height=height,
561
+ width=width,
562
+ num_inference_steps=steps,
563
+ guidance_scale=cfg_scale,
564
+ num_images_per_prompt=n,
565
+ generator=generator,
566
+ ).images
567
+ finally:
568
+ if args.backend == "glm" and hasattr(edit_pipeline, "vision_language_encoder"):
569
+ print("Moving GLM Vision Encoder back to CPU...")
570
+ edit_pipeline.vision_language_encoder.to("cpu")
571
+
572
+ for img in generated_images:
573
+ buffered = io.BytesIO()
574
+ img.save(buffered, format="PNG")
575
+ img_str = base64.b64encode(buffered.getvalue()).decode("utf-8")
576
+
577
+ if response_format == "b64_json":
578
+ response_images.append({"b64_json": img_str})
579
+ else:
580
+ # If url is requested we can't really do it without storage, so we fallback or error?
581
+ # For now, let's just assume simple b64_json as per request
582
+ response_images.append({"b64_json": img_str}) # Fallback
583
+
584
+ except Exception as e:
585
+ print(f"Error during editing: {e}")
586
+ print(traceback.format_exc())
587
+ raise HTTPException(status_code=500, detail=str(e))
588
+ finally:
589
+ flush()
590
+
591
+ t_req_total = time.perf_counter() - t_req_start
592
+ timings = timing_tracker.get_summary(t_req_total)
593
+ print(f"[TIMING] Text Encoder: {timings['text_encoder_s']:.2f}s | "
594
+ f"Transformer ({timings['steps']} steps): {timings['transformer_s']:.2f}s ({timings['step_avg_s']:.2f}s/step) | "
595
+ f"VAE: {timings['vae_s']:.2f}s | Total: {timings['total_s']:.2f}s")
596
+
597
+ return {
598
+ "created": int(time.time()),
599
+ "timings": timings,
600
+ "seed": seed,
601
+ "width": width,
602
+ "height": height,
603
+ "data": response_images
604
+ }
605
+
606
+
607
+
608
+ @app.post("/v1/images/generations")
609
+ async def generate_image(request: ImageGenerationRequest):
610
+ if not pipeline:
611
+ raise HTTPException(status_code=500, detail="Model not loaded")
612
+
613
+ if sleep_requested or is_sleeping_flag:
614
+ raise HTTPException(status_code=503, detail="Server is sleeping or trying to sleep.")
615
+
616
+ async with request_lock:
617
+ #print(f"Received generation request: {request.prompt}")
618
+
619
+ # Parse size
620
+ try:
621
+ width, height = map(int, request.size.split("x"))
622
+ except ValueError:
623
+ width, height = 1024, 1024
624
+
625
+ # Ensure dimensions are aligned to 32
626
+ width = (width // IMAGE_DIMENSION_ALIGNMENT) * IMAGE_DIMENSION_ALIGNMENT
627
+ height = (height // IMAGE_DIMENSION_ALIGNMENT) * IMAGE_DIMENSION_ALIGNMENT
628
+
629
+ response_images = []
630
+ timing_tracker.reset()
631
+ t_req_start = time.perf_counter()
632
+
633
+ try:
634
+ # Generate images (no image argument for txt2img!)
635
+ steps = request.num_inference_steps if request.num_inference_steps is not None else args.steps
636
+ cfg_scale = request.guidance_scale if request.guidance_scale is not None else args.guidance_scale
637
+ # negative_prompt not in standard request body in original snippet, but we added it to model
638
+ neg_prompt = request.negative_prompt if request.negative_prompt is not None else ""
639
+
640
+ generator = None
641
+ import random
642
+ seed = request.seed
643
+ if seed is None:
644
+ seed = random.randint(0, 2**32 - 1)
645
+
646
+ print(f"Using seed: {seed}")
647
+ generator = torch.Generator(device="cuda").manual_seed(seed)
648
+
649
+ if args.backend.startswith("qwen"):
650
+ if hasattr(pipeline, "transformer"):
651
+ set_zero_cond_t(pipeline.transformer, False)
652
+ generated_images = pipeline(
653
+ prompt=request.prompt,
654
+ height=height,
655
+ width=width,
656
+ num_inference_steps=steps,
657
+ true_cfg_scale=cfg_scale,
658
+ num_images_per_prompt=request.n,
659
+ negative_prompt=neg_prompt,
660
+ generator=generator,
661
+ ).images
662
+ else:
663
+ generated_images = pipeline(
664
+ prompt=request.prompt,
665
+ height=height,
666
+ width=width,
667
+ num_inference_steps=steps,
668
+ guidance_scale=cfg_scale,
669
+ num_images_per_prompt=request.n,
670
+ generator=generator,
671
+ # Not passing negative_prompt here for generation unless we confirm support in standard Flux pipeline?
672
+ ).images
673
+
674
+ for img in generated_images:
675
+ buffered = io.BytesIO()
676
+ img.save(buffered, format="PNG")
677
+ img_str = base64.b64encode(buffered.getvalue()).decode("utf-8")
678
+ response_images.append({"b64_json": img_str})
679
+
680
+ except Exception as e:
681
+ print(f"Error during generation: {e}")
682
+ print(traceback.format_exc())
683
+ raise HTTPException(status_code=500, detail=str(e))
684
+ finally:
685
+ flush()
686
+
687
+ t_req_total = time.perf_counter() - t_req_start
688
+ timings = timing_tracker.get_summary(t_req_total)
689
+ print(f"[TIMING] Text Encoder: {timings['text_encoder_s']:.2f}s | "
690
+ f"Transformer ({timings['steps']} steps): {timings['transformer_s']:.2f}s ({timings['step_avg_s']:.2f}s/step) | "
691
+ f"VAE: {timings['vae_s']:.2f}s | Total: {timings['total_s']:.2f}s")
692
+
693
+ return {
694
+ "created": int(time.time()),
695
+ "timings": timings,
696
+ "seed": seed,
697
+ "width": width,
698
+ "height": height,
699
+ "data": response_images
700
+ }
701
+
702
+ if __name__ == "__main__":
703
+ uvicorn.run(app, host=args.host, port=args.port)
extras/QwenImage21Backend.py ADDED
@@ -0,0 +1,73 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import torch
2
+ import os
3
+ from diffusers import QwenImage21Pipeline
4
+
5
+ # Channeling Nikola's cheerful seeker spirit!
6
+ # A clean, modular backend for Qwen-Image-2.1 with sequential layerwise offloading!
7
+
8
+ class QwenImage21Backend:
9
+ def __init__(
10
+ self,
11
+ model_id="/home/olegk/Nikola/models/Qwen/Qwen-Image-2.1",
12
+ gpu_id=0,
13
+ enable_tiling=True,
14
+ text_encoder_path=None,
15
+ ):
16
+ self.model_id = model_id
17
+ self.gpu_id = gpu_id
18
+ self.enable_tiling = enable_tiling
19
+ self.text_encoder_path = text_encoder_path
20
+ self.pipeline = None
21
+
22
+ def load(self):
23
+ print(f"Loading QwenImage21Backend from {self.model_id}...")
24
+ if self.text_encoder_path:
25
+ print(f" • Custom Text Encoder: {self.text_encoder_path}")
26
+ from transformers import Qwen3VLForConditionalGeneration, Qwen3VLProcessor
27
+
28
+ te_dir = (
29
+ os.path.join(self.text_encoder_path, "text_encoder")
30
+ if os.path.isdir(os.path.join(self.text_encoder_path, "text_encoder"))
31
+ else self.text_encoder_path
32
+ )
33
+ proc_dir = (
34
+ os.path.join(self.text_encoder_path, "processor")
35
+ if os.path.isdir(os.path.join(self.text_encoder_path, "processor"))
36
+ else self.text_encoder_path
37
+ )
38
+ if not os.path.exists(os.path.join(proc_dir, "tokenizer.json")):
39
+ proc_dir = os.path.join(self.model_id, "processor")
40
+
41
+ custom_te = Qwen3VLForConditionalGeneration.from_pretrained(
42
+ te_dir,
43
+ torch_dtype=torch.bfloat16,
44
+ low_cpu_mem_usage=True,
45
+ )
46
+ custom_proc = Qwen3VLProcessor.from_pretrained(proc_dir)
47
+
48
+ pipeline = QwenImage21Pipeline.from_pretrained(
49
+ self.model_id,
50
+ text_encoder=custom_te,
51
+ processor=custom_proc,
52
+ torch_dtype=torch.bfloat16,
53
+ )
54
+ else:
55
+ pipeline = QwenImage21Pipeline.from_pretrained(
56
+ self.model_id,
57
+ torch_dtype=torch.bfloat16,
58
+ )
59
+
60
+ print(f"Attaching sequential CPU offload on GPU {self.gpu_id} for layerwise execution...")
61
+ pipeline.enable_sequential_cpu_offload(gpu_id=self.gpu_id)
62
+
63
+ if self.enable_tiling:
64
+ print("Enabling VAE tiling for low-memory decode...")
65
+ try:
66
+ pipeline.vae.enable_tiling()
67
+ except Exception as e:
68
+ print(f"Note: VAE tiling could not be enabled ({e}), continuing with standard decode.")
69
+
70
+ self.pipeline = pipeline
71
+ # QwenImage21Pipeline natively handles both generations (t2i) and edits (i2i)
72
+ return self.pipeline, self.pipeline
73
+
extras/QwenImage21NVFP4Backend.py ADDED
@@ -0,0 +1,166 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # -*- coding: utf-8 -*-
2
+ """Nunchaku NVFP4 Resident Backend for Qwen-Image-2.1.
3
+
4
+ Loads the forged SVDQuant NVFP4 r32 Qwen-Image-2.1 transformer directly into VRAM,
5
+ enabling blazingly fast inference with zero layerwise PCIe streaming bottlenecks.
6
+ """
7
+
8
+ import os
9
+ import sys
10
+ import torch
11
+ from diffusers import QwenImage21Pipeline
12
+
13
+ # Ensure local packages are on path
14
+ ROOT_DIR = "/auto/home/amano/olegk/Nikola"
15
+ for p in [f"{ROOT_DIR}/packages/nunchaku", f"{ROOT_DIR}/packages/deepcompressor", ROOT_DIR]:
16
+ if p not in sys.path:
17
+ sys.path.insert(0, p)
18
+
19
+ from nunchaku.models.transformers.transformer_qwenimage21 import NunchakuQwenImage21Transformer2DModel
20
+
21
+
22
+ class QwenImage21NVFP4Backend:
23
+ def __init__(
24
+ self,
25
+ model_id="/home/olegk/Nikola/models/Qwen/Qwen-Image-2.1",
26
+ optimized_model_path="/home/olegk/Nikola/models/nunchaku-qwen-image-2.1/best_quality_fp4.safetensors",
27
+ gpu_id=0,
28
+ enable_tiling=True,
29
+ dynamic_scale_k=0.0,
30
+ stream_text_encoder=True,
31
+ text_encoder_path=None,
32
+ ):
33
+ self.model_id = model_id
34
+ self.optimized_model_path = optimized_model_path
35
+ self.gpu_id = gpu_id
36
+ self.enable_tiling = enable_tiling
37
+ self.dynamic_scale_k = dynamic_scale_k
38
+ self.stream_text_encoder = stream_text_encoder
39
+ self.text_encoder_path = text_encoder_path
40
+ self.pipeline = None
41
+
42
+ def load(self):
43
+ print(f"Loading QwenImage21NVFP4Backend...")
44
+ print(f" • Base Pipeline: {self.model_id}")
45
+ print(f" • Quantized DiT: {self.optimized_model_path}")
46
+ print(f" • Target GPU: cuda:{self.gpu_id}")
47
+ print(f" • Stream Text Encoder: {self.stream_text_encoder}")
48
+ if self.text_encoder_path:
49
+ print(f" • Custom Text Encoder: {self.text_encoder_path}")
50
+
51
+ device = f"cuda:{self.gpu_id}"
52
+
53
+ # 1. Load base pipeline without heavy transformer
54
+ # Pass transformer=None or dummy to save loading 13GB BF16 transformer
55
+ print("Loading peripheral pipeline components (Text Encoder, Tokenizer, VAE, Scheduler)...")
56
+ if self.text_encoder_path:
57
+ from transformers import Qwen3VLForConditionalGeneration, Qwen3VLProcessor
58
+
59
+ te_dir = (
60
+ os.path.join(self.text_encoder_path, "text_encoder")
61
+ if os.path.isdir(os.path.join(self.text_encoder_path, "text_encoder"))
62
+ else self.text_encoder_path
63
+ )
64
+ proc_dir = (
65
+ os.path.join(self.text_encoder_path, "processor")
66
+ if os.path.isdir(os.path.join(self.text_encoder_path, "processor"))
67
+ else self.text_encoder_path
68
+ )
69
+ if not os.path.exists(os.path.join(proc_dir, "tokenizer.json")):
70
+ proc_dir = os.path.join(self.model_id, "processor")
71
+
72
+ print(f" • Loading custom Qwen3-VL text encoder from: {te_dir}")
73
+ custom_te = Qwen3VLForConditionalGeneration.from_pretrained(
74
+ te_dir,
75
+ torch_dtype=torch.bfloat16,
76
+ low_cpu_mem_usage=True,
77
+ )
78
+ print(f" • Loading processor from: {proc_dir}")
79
+ custom_proc = Qwen3VLProcessor.from_pretrained(proc_dir)
80
+
81
+ pipeline = QwenImage21Pipeline.from_pretrained(
82
+ self.model_id,
83
+ transformer=None,
84
+ text_encoder=custom_te,
85
+ processor=custom_proc,
86
+ torch_dtype=torch.bfloat16,
87
+ )
88
+ else:
89
+ pipeline = QwenImage21Pipeline.from_pretrained(
90
+ self.model_id,
91
+ transformer=None,
92
+ torch_dtype=torch.bfloat16,
93
+ )
94
+
95
+ # 2. Load forged Nunchaku NVFP4 transformer directly into resident VRAM
96
+ print(f"Loading resident NVFP4 DiT into {device}...")
97
+ quantized_transformer = NunchakuQwenImage21Transformer2DModel.from_pretrained(
98
+ self.optimized_model_path,
99
+ device=device,
100
+ torch_dtype=torch.bfloat16,
101
+ )
102
+ pipeline.transformer = quantized_transformer
103
+
104
+ # 3. Place VAE directly on target device with memory-safe tiling
105
+ print(f"Placing VAE on {device}...")
106
+ pipeline.vae = pipeline.vae.to(device)
107
+ if self.enable_tiling:
108
+ try:
109
+ pipeline.vae.enable_tiling()
110
+ print("VAE tiling enabled successfully.")
111
+ except Exception as e:
112
+ print(f"Note: VAE tiling could not be enabled ({e})")
113
+
114
+ # 4. Text Encoder Configuration: Layerwise PCIe Streaming or CPU Fallback
115
+ if self.stream_text_encoder:
116
+ print(f"Configuring PCIe Layerwise Weight Streaming for Qwen3-VL on {device}...")
117
+ from stream_encoder import attach_qwen3vl_streamer
118
+ self.streamer = attach_qwen3vl_streamer(pipeline, device=device)
119
+
120
+ orig_encode_prompt = pipeline.encode_prompt
121
+
122
+ def streamed_safe_encode_prompt(*args, **kwargs):
123
+ kwargs.pop("device", None)
124
+ embeds = orig_encode_prompt(*args, device=torch.device(device), **kwargs)
125
+ target_dev = pipeline.transformer.device
126
+ return tuple(x.to(target_dev) if isinstance(x, torch.Tensor) else x for x in embeds)
127
+
128
+ pipeline.encode_prompt = streamed_safe_encode_prompt
129
+ else:
130
+ print(f"Configuring hybrid CPU prompt encoding (DiT & VAE 100% resident on {device})...")
131
+ orig_encode_prompt = pipeline.encode_prompt
132
+
133
+ def safe_encode_prompt(*args, **kwargs):
134
+ kwargs.pop("device", None)
135
+ te_device = pipeline.text_encoder.device
136
+ embeds = orig_encode_prompt(*args, device=te_device, **kwargs)
137
+ target_dev = pipeline.transformer.device
138
+ return tuple(x.to(target_dev) if isinstance(x, torch.Tensor) else x for x in embeds)
139
+
140
+ pipeline.encode_prompt = safe_encode_prompt
141
+
142
+ # 5. Automatically attach optimal dynamic latent variance damping s(t)
143
+ if self.dynamic_scale_k > 0.0:
144
+ k = self.dynamic_scale_k
145
+ print(f"Attaching automatic late-stage variance damping (k={k:.3f}, +1.20 dB fidelity boost)...")
146
+ orig_call = pipeline.__call__
147
+
148
+ def scaled_call(*args, **kwargs):
149
+ user_cb = kwargs.pop("callback_on_step_end", None)
150
+
151
+ def combined_cb(pipe_obj, step_idx, timestep, callback_kwargs):
152
+ t = float(timestep.item() if isinstance(timestep, torch.Tensor) else timestep)
153
+ if t < 600.0:
154
+ scale = 1.0 - k * ((600.0 - t) / 600.0)
155
+ callback_kwargs["latents"] = callback_kwargs["latents"] * scale
156
+ if user_cb is not None:
157
+ return user_cb(pipe_obj, step_idx, timestep, callback_kwargs)
158
+ return callback_kwargs
159
+
160
+ kwargs["callback_on_step_end"] = combined_cb
161
+ return orig_call(*args, **kwargs)
162
+
163
+ pipeline.__call__ = scaled_call
164
+
165
+ self.pipeline = pipeline
166
+ return self.pipeline, self.pipeline
extras/create_heretic_text_encoder.py ADDED
@@ -0,0 +1,201 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # -*- coding: utf-8 -*-
2
+ """Heretic Abliteration Engine for Qwen3-VL Text Encoder.
3
+
4
+ Applies Norm-Preserving Biprojected Abliteration (Heretic/Arditi et al.) to Qwen3-VL-8B:
5
+ 1. Measures the refusal/hesitation direction across all 36 transformer decoder layers.
6
+ 2. Orthogonalizes the refusal direction against the benign semantic direction.
7
+ 3. Applies rank-1 directional ablation to `self_attn.o_proj` and `mlp.down_proj`:
8
+ W' = normalize(W_norm - lambda * v * (v^T W_norm)) * ||W||_row
9
+ preserving exact row norms to protect general capability and language quality.
10
+ 4. Exports the decensored, hesitation-free text encoder to ~/Nikola/models/Qwen_Text_Encoder_Heretic.
11
+ """
12
+
13
+ import os
14
+ import sys
15
+ import shutil
16
+ import math
17
+ import torch
18
+ import torch.nn.functional as F
19
+ import torch.linalg as LA
20
+ from diffusers import QwenImage21Pipeline
21
+
22
+ ROOT_DIR = "/auto/home/amano/olegk/Nikola"
23
+ for p in [f"{ROOT_DIR}/src/imagegen", ROOT_DIR]:
24
+ if p not in sys.path:
25
+ sys.path.insert(0, p)
26
+
27
+ from stream_encoder import attach_qwen3vl_streamer
28
+
29
+ MODEL_PATH = "/home/olegk/Nikola/models/Qwen/Qwen-Image-2.1"
30
+ OUTPUT_DIR = "/home/olegk/Nikola/models/Qwen21_Text_Encoder_Heretic"
31
+
32
+ # Contrastive calibration prompts
33
+ BENIGN_PROMPTS = [
34
+ "A scenic view of a quiet tropical beach with palm trees, white sand, and gentle ocean waves under sunlight.",
35
+ "A group of athletic friends playing a casual game of volleyball on a sports court on a bright afternoon.",
36
+ "Two cheerful anime girls laughing together and talking in a colorful modern classroom setting.",
37
+ "A beautiful golden hour sunset over the coastal dunes with warm ambient lighting and ocean breeze.",
38
+ "Two athletic anime girls in standard sportswear practicing drills together on an outdoor gymnasium court.",
39
+ "A detailed photograph of a clean sandy beach with clear turquoise water and blue sky with light clouds.",
40
+ "Two anime girls taking a relaxing stroll along the seaside promenade during summer vacation.",
41
+ "A dynamic action photograph of a volleyball match on the sand with players in motion."
42
+ ]
43
+
44
+ SENSITIVE_PROMPTS = [
45
+ "Two anime girls in colorful skimpy bikinis playing beach volleyball on a sunny tropical beach.",
46
+ "Two attractive anime girls wearing revealing two-piece swimsuits posing playfully under the bright sun.",
47
+ "Two cute anime girls in skimpy micro bikinis jumping to spike a volleyball on the beach sand.",
48
+ "A beach volleyball match with two anime girls wearing small swimsuits, athletic dynamic poses.",
49
+ "Two anime girls in alluring colorful bikinis having fun playing sports on a sunny seaside beach.",
50
+ "A close up action shot of two anime girls in tiny bikinis diving for a volleyball on the beach.",
51
+ "Two beautiful anime girls wearing revealing beachwear and swimsuits posing by the ocean shoreline.",
52
+ "An alluring tropical beach scene with two anime girls in revealing bikinis playing athletic beach sports."
53
+ ]
54
+
55
+
56
+ def abliterate_layer_weights(weight: torch.Tensor, v: torch.Tensor, weight_factor: float) -> torch.Tensor:
57
+ """Norm-preserving biprojected abliteration: delta W = -lambda * v * (v^T W)."""
58
+ if weight_factor <= 0.0:
59
+ return weight
60
+
61
+ orig_dtype = weight.dtype
62
+ W = weight.float()
63
+ v = v.float().to(W.device)
64
+
65
+ # Calculate row norms
66
+ row_norms = LA.vector_norm(W, dim=1, keepdim=True)
67
+ # Normalize rows
68
+ W_norm = F.normalize(W, p=2, dim=1)
69
+
70
+ # v @ W_norm -> (in_features,)
71
+ lora_A = (v @ W_norm).view(1, -1)
72
+ # -weight_factor * v -> (out_features, 1)
73
+ lora_B = (-weight_factor * v).view(-1, 1)
74
+
75
+ # Project and renormalize
76
+ W_adj = W_norm + lora_B @ lora_A
77
+ W_adj = F.normalize(W_adj, p=2, dim=1)
78
+ W_final = W_adj * row_norms
79
+
80
+ return W_final.to(orig_dtype)
81
+
82
+
83
+ def create_heretic_model():
84
+ print("=" * 80)
85
+ print("🔮 FORGING OPTIMIZED HERETIC TEXT ENCODER (QWEN3-VL-8B)")
86
+ print("=" * 80)
87
+ print(f"Base Model: {MODEL_PATH}")
88
+ print(f"Output Directory: {OUTPUT_DIR}")
89
+
90
+ # 1. Load pipeline and attach layerwise streamer for rapid residual extraction
91
+ print("\n[Step 1/4] Loading pipeline & attaching layerwise streamer...")
92
+ pipe = QwenImage21Pipeline.from_pretrained(
93
+ MODEL_PATH,
94
+ transformer=None,
95
+ torch_dtype=torch.bfloat16,
96
+ )
97
+ streamer = attach_qwen3vl_streamer(pipe, device="cuda:0")
98
+
99
+ # 2. Extract layer-by-layer residuals across contrastive prompt sets
100
+ print("\n[Step 2/4] Measuring refusal directions across all 36 decoder layers...")
101
+
102
+ def get_mean_residuals(prompts, label):
103
+ layer_states = [[] for _ in range(37)]
104
+ print(f" • Collecting residuals for {len(prompts)} {label} prompts...")
105
+ for p in prompts:
106
+ fmt = pipe.prompt_template_t2i.format(p)
107
+ inputs = pipe.processor(text=[fmt], return_tensors="pt").to("cuda:0")
108
+ with torch.no_grad():
109
+ out = pipe.text_encoder(
110
+ input_ids=inputs.input_ids,
111
+ attention_mask=inputs.attention_mask,
112
+ output_hidden_states=True,
113
+ )
114
+ for l in range(37):
115
+ layer_states[l].append(out.hidden_states[l][0, -1, :].float().cpu())
116
+ return torch.stack([torch.stack(l).mean(dim=0) for l in layer_states])
117
+
118
+ b_means = get_mean_residuals(BENIGN_PROMPTS, "benign")
119
+ s_means = get_mean_residuals(SENSITIVE_PROMPTS, "sensitive")
120
+
121
+ # Compute difference of means
122
+ residual_directions = s_means - b_means
123
+
124
+ # Orthogonalize against benign direction (Heretic projected abliteration)
125
+ print(" • Orthogonalizing refusal directions against benign semantic vectors...")
126
+ good_directions = F.normalize(b_means, p=2, dim=1)
127
+ proj = torch.sum(residual_directions * good_directions, dim=1, keepdim=True)
128
+ ortho_directions = residual_directions - proj * good_directions
129
+ ortho_directions = F.normalize(ortho_directions, p=2, dim=1)
130
+
131
+ # 3. Apply Norm-Preserving Biprojected Abliteration to Language Model Weights
132
+ print("\n[Step 3/4] Applying Heretic norm-preserving abliteration to weights...")
133
+ lm = pipe.text_encoder.model.language_model
134
+ num_layers = len(lm.layers)
135
+
136
+ # Target layers: layers 16 to 35, centered at layer 26 with max_weight = 1.0
137
+ center_layer = 26.0
138
+ spread = 8.0
139
+
140
+ ablated_count = 0
141
+ for l_idx, layer in enumerate(lm.layers):
142
+ dist = abs(l_idx - center_layer)
143
+ # Smooth bell-shaped abliteration profile
144
+ weight_factor = float(math.exp(-(dist**2) / (2 * (spread**2))))
145
+ if weight_factor < 0.10:
146
+ weight_factor = 0.0 # skip early layers (0..12) where refusal is inactive
147
+
148
+ v = ortho_directions[l_idx + 1] # l_idx+1 accounts for embedding layer at idx 0
149
+
150
+ if weight_factor > 0.0:
151
+ print(f" • Layer {l_idx:2d}: Abliterating with lambda={weight_factor:.3f}...")
152
+
153
+ # 1. Attention Out Projection
154
+ if hasattr(layer.self_attn, "o_proj"):
155
+ with torch.no_grad():
156
+ layer.self_attn.o_proj.weight.data = abliterate_layer_weights(
157
+ layer.self_attn.o_proj.weight.data, v, weight_factor
158
+ )
159
+ ablated_count += 1
160
+
161
+ # 2. MLP Down Projection
162
+ if hasattr(layer.mlp, "down_proj"):
163
+ with torch.no_grad():
164
+ layer.mlp.down_proj.weight.data = abliterate_layer_weights(
165
+ layer.mlp.down_proj.weight.data, v, weight_factor
166
+ )
167
+ ablated_count += 1
168
+
169
+ print(f" • Successfully abliterated {ablated_count} linear projection matrices!")
170
+
171
+ # 4. Save Standalone Checkpoint to models/Qwen_Text_Encoder_Heretic
172
+ print(f"\n[Step 4/4] Saving Heretic Text Encoder to {OUTPUT_DIR}...")
173
+ os.makedirs(OUTPUT_DIR, exist_ok=True)
174
+
175
+ # Save text_encoder subfolder (for direct use with Diffusers or Transformers)
176
+ sub_dir = os.path.join(OUTPUT_DIR, "text_encoder")
177
+ os.makedirs(sub_dir, exist_ok=True)
178
+ pipe.text_encoder.save_pretrained(sub_dir)
179
+ print(f" • Saved abliterated text encoder weights to: {sub_dir}")
180
+
181
+ # Also copy tokenizer and processor files
182
+ proc_src = os.path.join(MODEL_PATH, "processor")
183
+ proc_dst = os.path.join(OUTPUT_DIR, "processor")
184
+ if os.path.exists(proc_src):
185
+ if os.path.exists(proc_dst):
186
+ shutil.rmtree(proc_dst)
187
+ shutil.copytree(proc_src, proc_dst)
188
+ print(f" • Copied processor to: {proc_dst}")
189
+
190
+ # Copy root config files
191
+ for fname in ["config.json", "generation_config.json"]:
192
+ src_f = os.path.join(MODEL_PATH, "text_encoder", fname)
193
+ if os.path.exists(src_f):
194
+ shutil.copy(src_f, os.path.join(OUTPUT_DIR, fname))
195
+
196
+ print("\n🎉 SUCCESS: Heretic Text Encoder forged and saved successfully!")
197
+ print("=" * 80)
198
+
199
+
200
+ if __name__ == "__main__":
201
+ create_heretic_model()
extras/imagegen_qwen21_nvfp4.sh ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/bin/bash
2
+ # Start ImageEditServer with forged SVDQuant NVFP4 r32 Qwen-Image-2.1 model
3
+ # Runs 100% resident in VRAM on GPU 0 with Heretic text encoder by default
4
+ # Default: 25 steps, guidance-scale 1.0, port 4500, Text Encoder: Qwen21_Text_Encoder_Heretic
5
+
6
+ PORT=${1:-4500}
7
+ MODEL_DIR="/home/olegk/Nikola/models/Qwen/Qwen-Image-2.1"
8
+ OPTIMIZED_MODEL="/home/olegk/Nikola/models/nunchaku-qwen-image-2.1/best_quality_fp4.safetensors"
9
+ TEXT_ENCODER=${2:-"/home/olegk/Nikola/models/Qwen21_Text_Encoder_Heretic"}
10
+
11
+ echo "=========================================================="
12
+ echo "Starting ImageEditServer (Qwen-Image-2.1 SVDQuant NVFP4 - Resident VRAM)"
13
+ echo "Port: $PORT"
14
+ echo "Model: $MODEL_DIR"
15
+ echo "Optimized Model: $OPTIMIZED_MODEL"
16
+ echo "Text Encoder: $TEXT_ENCODER"
17
+ echo "Steps: 25 | Guidance Scale: 1.0 | Resident Mode: True"
18
+ echo "=========================================================="
19
+
20
+ export CUDA_VISIBLE_DEVICES="0"
21
+ export PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"
22
+ export PYTHONPATH="/home/olegk/Nikola/packages/nunchaku:/home/olegk/Nikola/packages/deepcompressor:/home/olegk/Nikola:$PYTHONPATH"
23
+
24
+ cd /auto/home/amano/olegk/Nikola/src/imagegen
25
+
26
+ EXTRA_ARGS=""
27
+ if [ -n "$TEXT_ENCODER" ] && [ "$TEXT_ENCODER" != "none" ]; then
28
+ EXTRA_ARGS="--text-encoder $TEXT_ENCODER"
29
+ fi
30
+
31
+ exec /home/olegk/Nikola/.venv/bin/python ImageEditServer.py \
32
+ --host 0.0.0.0 \
33
+ --port "$PORT" \
34
+ --model "$MODEL_DIR" \
35
+ --optimized-model "$OPTIMIZED_MODEL" \
36
+ --backend qwen21-nvfp4 \
37
+ --steps 25 \
38
+ --guidance-scale 1.0 \
39
+ $EXTRA_ARGS
40
+
extras/stream_encoder.py ADDED
@@ -0,0 +1,190 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # -*- coding: utf-8 -*-
2
+ """Layerwise Weight Streaming Engine for Qwen3-VL-8B in Qwen-Image-2.1.
3
+
4
+ Pins all 36 language model decoder layers in host RAM and streams them through a single
5
+ pre-allocated GPU layer buffer (368 MB VRAM) over PCIe (~28.7 GB/s).
6
+ Keeps the visual ViT encoder (1.07 GB) resident on GPU.
7
+
8
+ Achieves GPU compute speeds (1.06s multimodal prompt encode vs 17.18s on CPU, saving >16s per edit)
9
+ with only ~1.45 GB VRAM footprint and 100% bit-exact mathematical parity (zero quality loss).
10
+ """
11
+
12
+ import copy
13
+ import time
14
+ import torch
15
+ import torch.nn as nn
16
+ from transformers.modeling_outputs import BaseModelOutputWithPast
17
+ from transformers.models.qwen3_vl.modeling_qwen3_vl import create_causal_mask
18
+
19
+
20
+ class Qwen3VLLayerwiseStreamer:
21
+ """Streams Qwen3-VL language model decoder layers through a single static GPU buffer."""
22
+
23
+ def __init__(self, pipeline, device="cuda:0"):
24
+ self.pipeline = pipeline
25
+ self.device = torch.device(device)
26
+ self.text_encoder = pipeline.text_encoder
27
+ self.lm = getattr(self.text_encoder.model, "language_model", self.text_encoder.model)
28
+ self.num_layers = len(self.lm.layers)
29
+
30
+ print(f"Initializing Qwen3VLLayerwiseStreamer for {self.num_layers} layers on {self.device}...")
31
+
32
+ # 1. Pin CPU layers in host memory for maximum PCIe transfer throughput
33
+ t0 = time.perf_counter()
34
+ self.cpu_layers = []
35
+ for layer in self.lm.layers:
36
+ layer = layer.to("cpu", dtype=torch.bfloat16)
37
+ for p in layer.parameters():
38
+ if not p.data.is_pinned():
39
+ p.data = p.data.pin_memory()
40
+ for b in layer.buffers():
41
+ if not b.data.is_pinned():
42
+ b.data = b.data.pin_memory()
43
+ self.cpu_layers.append(layer)
44
+ t_pin = time.perf_counter() - t0
45
+ print(f" • Pinned {self.num_layers} layers in CPU RAM in {t_pin:.2f} s")
46
+
47
+ # 2. Allocate ONE single template GPU layer buffer in VRAM (~368 MB)
48
+ self.gpu_layer = copy.deepcopy(self.cpu_layers[0]).to(self.device, dtype=torch.bfloat16)
49
+ gpu_param_dict = dict(self.gpu_layer.named_parameters())
50
+ gpu_buffer_dict = dict(self.gpu_layer.named_buffers())
51
+
52
+ # Pre-build parameter transfer pairs for zero-overhead non-blocking copying
53
+ self.param_pairs = []
54
+ for i in range(self.num_layers):
55
+ lp = [(gpu_param_dict[name], cp) for name, cp in self.cpu_layers[i].named_parameters()]
56
+ lb = [(gpu_buffer_dict[name], cb) for name, cb in self.cpu_layers[i].named_buffers()]
57
+ self.param_pairs.append((lp, lb))
58
+
59
+ gpu_mb = sum(p.numel() * p.element_size() for p in self.gpu_layer.parameters()) / (1024**2)
60
+ print(f" • Static GPU layer buffer allocated: {gpu_mb:.2f} MB VRAM")
61
+
62
+ # 3. Place small peripheral layers directly on target GPU
63
+ self.lm.rotary_emb = self.lm.rotary_emb.to(self.device)
64
+ self.lm.embed_tokens = self.lm.embed_tokens.to(self.device)
65
+ self.lm.norm = self.lm.norm.to(self.device)
66
+
67
+ # 4. Place visual ViT encoder directly on target GPU (1.07 GB VRAM)
68
+ if hasattr(self.text_encoder.model, "visual") and self.text_encoder.model.visual is not None:
69
+ self.text_encoder.model.visual = self.text_encoder.model.visual.to(self.device, dtype=torch.bfloat16)
70
+ print(" • Visual ViT encoder placed resident on GPU (1.07 GB VRAM)")
71
+
72
+ # 5. Bypass unused lm_head (152,064 vocab projection, saving 1.24 GB computation)
73
+ class DummyHead(nn.Module):
74
+ def forward(self, x):
75
+ return None
76
+
77
+ self.text_encoder.lm_head = DummyHead()
78
+ print(" • Bypassed unused lm_head projection")
79
+
80
+ # 6. Install hooked forward pass
81
+ self.orig_lm_forward = self.lm.forward
82
+ self.lm.forward = self.streamed_forward
83
+
84
+ # 7. Route pipeline._get_qwen_prompt_embeds to target GPU
85
+ self.orig_get_embeds = self.pipeline._get_qwen_prompt_embeds
86
+ target_dev = self.device
87
+
88
+ def gpu_get_embeds(prompt_arg, image_arg, device_arg=None):
89
+ return self.orig_get_embeds(prompt_arg, image_arg, device=target_dev)
90
+
91
+ self.pipeline._get_qwen_prompt_embeds = gpu_get_embeds
92
+ print(f" • Hooked Qwen3-VL language model and prompt embedding router onto {self.device}!")
93
+
94
+ @torch.no_grad()
95
+ def streamed_forward(
96
+ self,
97
+ input_ids: torch.LongTensor | None = None,
98
+ attention_mask: torch.Tensor | None = None,
99
+ position_ids: torch.LongTensor | None = None,
100
+ past_key_values=None,
101
+ inputs_embeds: torch.FloatTensor | None = None,
102
+ use_cache: bool | None = None,
103
+ visual_pos_masks: torch.Tensor | None = None,
104
+ deepstack_visual_embeds: list[torch.Tensor] | None = None,
105
+ output_hidden_states: bool | None = None,
106
+ **kwargs,
107
+ ) -> BaseModelOutputWithPast:
108
+ """Executes language model decoding by streaming layers one by one into the GPU buffer."""
109
+ if inputs_embeds is None:
110
+ inputs_embeds = self.lm.embed_tokens(input_ids)
111
+ inputs_embeds = inputs_embeds.to(self.device)
112
+
113
+ if position_ids is None:
114
+ past_seen = past_key_values.get_seq_length() if past_key_values is not None else 0
115
+ position_ids = torch.arange(inputs_embeds.shape[1], device=self.device) + past_seen
116
+ position_ids = position_ids.view(1, 1, -1).expand(4, inputs_embeds.shape[0], -1)
117
+ elif position_ids.ndim == 2:
118
+ position_ids = position_ids[None, ...].expand(4, position_ids.shape[0], -1)
119
+
120
+ position_ids = position_ids.to(self.device)
121
+ if position_ids.ndim == 3 and position_ids.shape[0] == 4:
122
+ text_position_ids = position_ids[0]
123
+ rotary_pos_ids = position_ids[1:]
124
+ else:
125
+ text_position_ids = None
126
+ rotary_pos_ids = position_ids
127
+
128
+ causal_mask = create_causal_mask(
129
+ config=self.lm.config,
130
+ inputs_embeds=inputs_embeds,
131
+ attention_mask=attention_mask.to(self.device) if attention_mask is not None else None,
132
+ past_key_values=past_key_values,
133
+ position_ids=text_position_ids,
134
+ )
135
+
136
+ position_embeddings = self.lm.rotary_emb(inputs_embeds, rotary_pos_ids)
137
+ hidden_states = inputs_embeds
138
+
139
+ if visual_pos_masks is not None:
140
+ visual_pos_masks = visual_pos_masks.to(self.device)
141
+ if deepstack_visual_embeds is not None:
142
+ deepstack_visual_embeds = [d.to(self.device) for d in deepstack_visual_embeds]
143
+
144
+ all_hidden_states = () if output_hidden_states else None
145
+
146
+ # Stream all 36 decoder layers through the static GPU buffer
147
+ for layer_idx in range(self.num_layers):
148
+ if output_hidden_states:
149
+ all_hidden_states = all_hidden_states + (hidden_states,)
150
+
151
+ params, buffers = self.param_pairs[layer_idx]
152
+ for gp, cp in params:
153
+ gp.data.copy_(cp.data, non_blocking=True)
154
+ for gb, cb in buffers:
155
+ gb.data.copy_(cb.data, non_blocking=True)
156
+
157
+ layer_outputs = self.gpu_layer(
158
+ hidden_states,
159
+ attention_mask=causal_mask,
160
+ position_ids=text_position_ids,
161
+ past_key_values=past_key_values,
162
+ position_embeddings=position_embeddings,
163
+ **kwargs,
164
+ )
165
+ hidden_states = layer_outputs
166
+
167
+ # Add multi-layer deepstack visual features if present
168
+ if deepstack_visual_embeds is not None and layer_idx in range(len(deepstack_visual_embeds)):
169
+ hidden_states = self.lm._deepstack_process(
170
+ hidden_states,
171
+ visual_pos_masks,
172
+ deepstack_visual_embeds[layer_idx],
173
+ )
174
+
175
+ pre_norm_states = hidden_states
176
+ if output_hidden_states:
177
+ all_hidden_states = all_hidden_states + (pre_norm_states,)
178
+
179
+ norm_states = self.lm.norm(hidden_states)
180
+
181
+ return BaseModelOutputWithPast(
182
+ last_hidden_state=norm_states,
183
+ past_key_values=past_key_values,
184
+ hidden_states=all_hidden_states,
185
+ )
186
+
187
+
188
+ def attach_qwen3vl_streamer(pipeline, device="cuda:0") -> Qwen3VLLayerwiseStreamer:
189
+ """Convenience factory to attach layerwise streaming to any QwenImage21Pipeline."""
190
+ return Qwen3VLLayerwiseStreamer(pipeline, device=device)
extras/test_heretic_beach_volleyball.py ADDED
@@ -0,0 +1,128 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # -*- coding: utf-8 -*-
2
+ """Verification & Comparison of Heretic Decensored Text Encoder on Qwen-Image-2.1.
3
+
4
+ Generates identical text-to-image scenes comparing:
5
+ 1. Base / Stock Qwen3-VL Text Encoder (Safety Aligned / Blush Vector)
6
+ 2. Heretic Qwen3-VL Text Encoder (Abliterated / Hesitation-Free)
7
+
8
+ Target Prompt:
9
+ "Two cute anime girls in colorful bikinis playing beach volleyball on a sunny tropical beach,
10
+ dynamic action pose, jumping to spike the ball, sharp focus, vibrant anime style,
11
+ detailed background with ocean and palm trees"
12
+ """
13
+
14
+ import os
15
+ import sys
16
+ import time
17
+ import shutil
18
+ import torch
19
+ from PIL import Image
20
+
21
+ ROOT_DIR = "/auto/home/amano/olegk/Nikola"
22
+ for p in [f"{ROOT_DIR}/packages/nunchaku", f"{ROOT_DIR}/packages/deepcompressor", f"{ROOT_DIR}/src/imagegen", ROOT_DIR]:
23
+ if p not in sys.path:
24
+ sys.path.insert(0, p)
25
+
26
+ from QwenImage21NVFP4Backend import QwenImage21NVFP4Backend
27
+
28
+ PROMPT = (
29
+ "Two cute anime girls in colorful bikinis playing beach volleyball on a sunny tropical beach, "
30
+ "dynamic action pose, jumping to spike the ball, sharp focus, vibrant anime style, "
31
+ "detailed background with ocean and palm trees"
32
+ )
33
+ SEED = 42
34
+ STEPS = 25
35
+ GUIDANCE_SCALE = 1.0
36
+ HEIGHT = 1024
37
+ WIDTH = 1024
38
+
39
+ TMP_DIR = "/home/olegk/tmp"
40
+ ARTIFACT_DIR = "/home/olegk/.gemini/antigravity/brain/19d2cb51-1ef6-42f5-a95f-f814fe6ba720"
41
+ HERETIC_PATH = "/home/olegk/Nikola/models/Qwen21_Text_Encoder_Heretic"
42
+
43
+
44
+ def generate_with_backend(backend_name: str, text_encoder_path=None):
45
+ print(f"\n{'='*70}")
46
+ print(f"🚀 RUNNING GENERATION: {backend_name}")
47
+ print(f"Text Encoder Path: {text_encoder_path or 'Stock (Base Qwen-Image-2.1)'}")
48
+ print(f"{'='*70}")
49
+
50
+ backend = QwenImage21NVFP4Backend(
51
+ model_id="/home/olegk/Nikola/models/Qwen/Qwen-Image-2.1",
52
+ optimized_model_path="/home/olegk/Nikola/models/nunchaku-qwen-image-2.1/best_quality_fp4.safetensors",
53
+ gpu_id=0,
54
+ enable_tiling=True,
55
+ dynamic_scale_k=0.0,
56
+ stream_text_encoder=True,
57
+ text_encoder_path=text_encoder_path,
58
+ )
59
+
60
+ t0 = time.perf_counter()
61
+ pipeline, _ = backend.load()
62
+ t_load = time.perf_counter() - t0
63
+ print(f"Backend loaded in {t_load:.2f} s")
64
+
65
+ generator = torch.Generator(device="cuda:0").manual_seed(SEED)
66
+
67
+ torch.cuda.synchronize()
68
+ t_gen_start = time.perf_counter()
69
+ output = pipeline(
70
+ prompt=PROMPT,
71
+ height=HEIGHT,
72
+ width=WIDTH,
73
+ num_inference_steps=STEPS,
74
+ true_cfg_scale=GUIDANCE_SCALE,
75
+ generator=generator,
76
+ )
77
+ torch.cuda.synchronize()
78
+ t_gen = time.perf_counter() - t_gen_start
79
+ print(f"Generation completed in {t_gen:.2f} s")
80
+
81
+ img = output.images[0]
82
+
83
+ # Clean up pipeline from GPU memory
84
+ del pipeline
85
+ del backend
86
+ torch.cuda.empty_cache()
87
+
88
+ return img, t_gen
89
+
90
+
91
+ def main():
92
+ os.makedirs(TMP_DIR, exist_ok=True)
93
+ os.makedirs(ARTIFACT_DIR, exist_ok=True)
94
+
95
+ print("=" * 80)
96
+ print("🏐 ANIME BEACH VOLLEYBALL TEXT ENCODER COMPARISON TEST")
97
+ print(f"Prompt: {PROMPT}")
98
+ print(f"Seed: {SEED} | Steps: {STEPS} | CFG: {GUIDANCE_SCALE}")
99
+ print("=" * 80)
100
+
101
+ # 1. Run Stock Text Encoder
102
+ img_stock, t_stock = generate_with_backend("STOCK (Base Qwen-Image-2.1 Text Encoder)", text_encoder_path=None)
103
+ stock_path = os.path.join(TMP_DIR, "qwen21_stock_beach_volleyball.png")
104
+ stock_art = os.path.join(ARTIFACT_DIR, "qwen21_stock_beach_volleyball.png")
105
+ img_stock.save(stock_path)
106
+ shutil.copy(stock_path, stock_art)
107
+ print(f"Saved stock image to: {stock_path} and artifact")
108
+
109
+ # 2. Run Heretic Text Encoder
110
+ img_heretic, t_heretic = generate_with_backend(
111
+ "HERETIC (Decensored Qwen3-VL Text Encoder)",
112
+ text_encoder_path=HERETIC_PATH,
113
+ )
114
+ heretic_path = os.path.join(TMP_DIR, "qwen21_heretic_beach_volleyball.png")
115
+ heretic_art = os.path.join(ARTIFACT_DIR, "qwen21_heretic_beach_volleyball.png")
116
+ img_heretic.save(heretic_path)
117
+ shutil.copy(heretic_path, heretic_art)
118
+ print(f"Saved heretic image to: {heretic_path} and artifact")
119
+
120
+ print("\n" + "=" * 80)
121
+ print("🎉 COMPARISON TEST COMPLETED SUCCESSFULLY!")
122
+ print(f"Stock Image: {stock_path} ({t_stock:.2f} s)")
123
+ print(f"Heretic Image: {heretic_path} ({t_heretic:.2f} s)")
124
+ print("=" * 80)
125
+
126
+
127
+ if __name__ == "__main__":
128
+ main()
generation_config.json ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 151643,
3
+ "do_sample": true,
4
+ "eos_token_id": [
5
+ 151645,
6
+ 151643
7
+ ],
8
+ "pad_token_id": 151643,
9
+ "temperature": 0.7,
10
+ "top_k": 20,
11
+ "top_p": 0.8,
12
+ "transformers_version": "5.3.0.dev0"
13
+ }
merges.txt ADDED
The diff for this file is too large to render. See raw diff
 
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ba4cadb5ccea2f11dcbb7facce01cb35a76c9a2d5553b34f681b2e16227a1f02
3
+ size 16289680712
preprocessor_config.json ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "crop_size": null,
3
+ "data_format": "channels_first",
4
+ "default_to_square": true,
5
+ "device": null,
6
+ "disable_grouping": null,
7
+ "do_center_crop": null,
8
+ "do_convert_rgb": true,
9
+ "do_normalize": true,
10
+ "do_pad": null,
11
+ "do_rescale": true,
12
+ "do_resize": true,
13
+ "image_mean": [
14
+ 0.5,
15
+ 0.5,
16
+ 0.5
17
+ ],
18
+ "image_processor_type": "Qwen2VLImageProcessorFast",
19
+ "image_std": [
20
+ 0.5,
21
+ 0.5,
22
+ 0.5
23
+ ],
24
+ "input_data_format": null,
25
+ "max_pixels": null,
26
+ "merge_size": 2,
27
+ "min_pixels": null,
28
+ "pad_size": null,
29
+ "patch_size": 16,
30
+ "processor_class": "Qwen3VLProcessor",
31
+ "resample": 3,
32
+ "rescale_factor": 0.00392156862745098,
33
+ "return_tensors": null,
34
+ "size": {
35
+ "longest_edge": 16777216,
36
+ "shortest_edge": 65536
37
+ },
38
+ "temporal_patch_size": 2
39
+ }
processor/added_tokens.json ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "</think>": 151668,
3
+ "</tool_call>": 151658,
4
+ "</tool_response>": 151666,
5
+ "<think>": 151667,
6
+ "<tool_call>": 151657,
7
+ "<tool_response>": 151665,
8
+ "<|box_end|>": 151649,
9
+ "<|box_start|>": 151648,
10
+ "<|endoftext|>": 151643,
11
+ "<|file_sep|>": 151664,
12
+ "<|fim_middle|>": 151660,
13
+ "<|fim_pad|>": 151662,
14
+ "<|fim_prefix|>": 151659,
15
+ "<|fim_suffix|>": 151661,
16
+ "<|im_end|>": 151645,
17
+ "<|im_start|>": 151644,
18
+ "<|image_pad|>": 151655,
19
+ "<|object_ref_end|>": 151647,
20
+ "<|object_ref_start|>": 151646,
21
+ "<|quad_end|>": 151651,
22
+ "<|quad_start|>": 151650,
23
+ "<|repo_name|>": 151663,
24
+ "<|video_pad|>": 151656,
25
+ "<|vision_end|>": 151653,
26
+ "<|vision_pad|>": 151654,
27
+ "<|vision_start|>": 151652
28
+ }
processor/chat_template.jinja ADDED
@@ -0,0 +1,120 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- if tools %}
2
+ {{- '<|im_start|>system\n' }}
3
+ {%- if messages[0].role == 'system' %}
4
+ {%- if messages[0].content is string %}
5
+ {{- messages[0].content }}
6
+ {%- else %}
7
+ {%- for content in messages[0].content %}
8
+ {%- if 'text' in content %}
9
+ {{- content.text }}
10
+ {%- endif %}
11
+ {%- endfor %}
12
+ {%- endif %}
13
+ {{- '\n\n' }}
14
+ {%- endif %}
15
+ {{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
16
+ {%- for tool in tools %}
17
+ {{- "\n" }}
18
+ {{- tool | tojson }}
19
+ {%- endfor %}
20
+ {{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
21
+ {%- else %}
22
+ {%- if messages[0].role == 'system' %}
23
+ {{- '<|im_start|>system\n' }}
24
+ {%- if messages[0].content is string %}
25
+ {{- messages[0].content }}
26
+ {%- else %}
27
+ {%- for content in messages[0].content %}
28
+ {%- if 'text' in content %}
29
+ {{- content.text }}
30
+ {%- endif %}
31
+ {%- endfor %}
32
+ {%- endif %}
33
+ {{- '<|im_end|>\n' }}
34
+ {%- endif %}
35
+ {%- endif %}
36
+ {%- set image_count = namespace(value=0) %}
37
+ {%- set video_count = namespace(value=0) %}
38
+ {%- for message in messages %}
39
+ {%- if message.role == "user" %}
40
+ {{- '<|im_start|>' + message.role + '\n' }}
41
+ {%- if message.content is string %}
42
+ {{- message.content }}
43
+ {%- else %}
44
+ {%- for content in message.content %}
45
+ {%- if content.type == 'image' or 'image' in content or 'image_url' in content %}
46
+ {%- set image_count.value = image_count.value + 1 %}
47
+ {%- if add_vision_id %}Picture {{ image_count.value }}: {% endif -%}
48
+ <|vision_start|><|image_pad|><|vision_end|>
49
+ {%- elif content.type == 'video' or 'video' in content %}
50
+ {%- set video_count.value = video_count.value + 1 %}
51
+ {%- if add_vision_id %}Video {{ video_count.value }}: {% endif -%}
52
+ <|vision_start|><|video_pad|><|vision_end|>
53
+ {%- elif 'text' in content %}
54
+ {{- content.text }}
55
+ {%- endif %}
56
+ {%- endfor %}
57
+ {%- endif %}
58
+ {{- '<|im_end|>\n' }}
59
+ {%- elif message.role == "assistant" %}
60
+ {{- '<|im_start|>' + message.role + '\n' }}
61
+ {%- if message.content is string %}
62
+ {{- message.content }}
63
+ {%- else %}
64
+ {%- for content_item in message.content %}
65
+ {%- if 'text' in content_item %}
66
+ {{- content_item.text }}
67
+ {%- endif %}
68
+ {%- endfor %}
69
+ {%- endif %}
70
+ {%- if message.tool_calls %}
71
+ {%- for tool_call in message.tool_calls %}
72
+ {%- if (loop.first and message.content) or (not loop.first) %}
73
+ {{- '\n' }}
74
+ {%- endif %}
75
+ {%- if tool_call.function %}
76
+ {%- set tool_call = tool_call.function %}
77
+ {%- endif %}
78
+ {{- '<tool_call>\n{"name": "' }}
79
+ {{- tool_call.name }}
80
+ {{- '", "arguments": ' }}
81
+ {%- if tool_call.arguments is string %}
82
+ {{- tool_call.arguments }}
83
+ {%- else %}
84
+ {{- tool_call.arguments | tojson }}
85
+ {%- endif %}
86
+ {{- '}\n</tool_call>' }}
87
+ {%- endfor %}
88
+ {%- endif %}
89
+ {{- '<|im_end|>\n' }}
90
+ {%- elif message.role == "tool" %}
91
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
92
+ {{- '<|im_start|>user' }}
93
+ {%- endif %}
94
+ {{- '\n<tool_response>\n' }}
95
+ {%- if message.content is string %}
96
+ {{- message.content }}
97
+ {%- else %}
98
+ {%- for content in message.content %}
99
+ {%- if content.type == 'image' or 'image' in content or 'image_url' in content %}
100
+ {%- set image_count.value = image_count.value + 1 %}
101
+ {%- if add_vision_id %}Picture {{ image_count.value }}: {% endif -%}
102
+ <|vision_start|><|image_pad|><|vision_end|>
103
+ {%- elif content.type == 'video' or 'video' in content %}
104
+ {%- set video_count.value = video_count.value + 1 %}
105
+ {%- if add_vision_id %}Video {{ video_count.value }}: {% endif -%}
106
+ <|vision_start|><|video_pad|><|vision_end|>
107
+ {%- elif 'text' in content %}
108
+ {{- content.text }}
109
+ {%- endif %}
110
+ {%- endfor %}
111
+ {%- endif %}
112
+ {{- '\n</tool_response>' }}
113
+ {%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
114
+ {{- '<|im_end|>\n' }}
115
+ {%- endif %}
116
+ {%- endif %}
117
+ {%- endfor %}
118
+ {%- if add_generation_prompt %}
119
+ {{- '<|im_start|>assistant\n' }}
120
+ {%- endif %}
processor/merges.txt ADDED
The diff for this file is too large to render. See raw diff
 
processor/preprocessor_config.json ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "crop_size": null,
3
+ "data_format": "channels_first",
4
+ "default_to_square": true,
5
+ "device": null,
6
+ "disable_grouping": null,
7
+ "do_center_crop": null,
8
+ "do_convert_rgb": true,
9
+ "do_normalize": true,
10
+ "do_pad": null,
11
+ "do_rescale": true,
12
+ "do_resize": true,
13
+ "image_mean": [
14
+ 0.5,
15
+ 0.5,
16
+ 0.5
17
+ ],
18
+ "image_processor_type": "Qwen2VLImageProcessorFast",
19
+ "image_std": [
20
+ 0.5,
21
+ 0.5,
22
+ 0.5
23
+ ],
24
+ "input_data_format": null,
25
+ "max_pixels": null,
26
+ "merge_size": 2,
27
+ "min_pixels": null,
28
+ "pad_size": null,
29
+ "patch_size": 16,
30
+ "processor_class": "Qwen3VLProcessor",
31
+ "resample": 3,
32
+ "rescale_factor": 0.00392156862745098,
33
+ "return_tensors": null,
34
+ "size": {
35
+ "longest_edge": 16777216,
36
+ "shortest_edge": 65536
37
+ },
38
+ "temporal_patch_size": 2
39
+ }
processor/special_tokens_map.json ADDED
@@ -0,0 +1,31 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "additional_special_tokens": [
3
+ "<|im_start|>",
4
+ "<|im_end|>",
5
+ "<|object_ref_start|>",
6
+ "<|object_ref_end|>",
7
+ "<|box_start|>",
8
+ "<|box_end|>",
9
+ "<|quad_start|>",
10
+ "<|quad_end|>",
11
+ "<|vision_start|>",
12
+ "<|vision_end|>",
13
+ "<|vision_pad|>",
14
+ "<|image_pad|>",
15
+ "<|video_pad|>"
16
+ ],
17
+ "eos_token": {
18
+ "content": "<|im_end|>",
19
+ "lstrip": false,
20
+ "normalized": false,
21
+ "rstrip": false,
22
+ "single_word": false
23
+ },
24
+ "pad_token": {
25
+ "content": "<|endoftext|>",
26
+ "lstrip": false,
27
+ "normalized": false,
28
+ "rstrip": false,
29
+ "single_word": false
30
+ }
31
+ }
processor/tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:aeb13307a71acd8fe81861d94ad54ab689df773318809eed3cbe794b4492dae4
3
+ size 11422654
processor/tokenizer_config.json ADDED
@@ -0,0 +1,240 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_bos_token": false,
3
+ "add_prefix_space": false,
4
+ "added_tokens_decoder": {
5
+ "151643": {
6
+ "content": "<|endoftext|>",
7
+ "lstrip": false,
8
+ "normalized": false,
9
+ "rstrip": false,
10
+ "single_word": false,
11
+ "special": true
12
+ },
13
+ "151644": {
14
+ "content": "<|im_start|>",
15
+ "lstrip": false,
16
+ "normalized": false,
17
+ "rstrip": false,
18
+ "single_word": false,
19
+ "special": true
20
+ },
21
+ "151645": {
22
+ "content": "<|im_end|>",
23
+ "lstrip": false,
24
+ "normalized": false,
25
+ "rstrip": false,
26
+ "single_word": false,
27
+ "special": true
28
+ },
29
+ "151646": {
30
+ "content": "<|object_ref_start|>",
31
+ "lstrip": false,
32
+ "normalized": false,
33
+ "rstrip": false,
34
+ "single_word": false,
35
+ "special": true
36
+ },
37
+ "151647": {
38
+ "content": "<|object_ref_end|>",
39
+ "lstrip": false,
40
+ "normalized": false,
41
+ "rstrip": false,
42
+ "single_word": false,
43
+ "special": true
44
+ },
45
+ "151648": {
46
+ "content": "<|box_start|>",
47
+ "lstrip": false,
48
+ "normalized": false,
49
+ "rstrip": false,
50
+ "single_word": false,
51
+ "special": true
52
+ },
53
+ "151649": {
54
+ "content": "<|box_end|>",
55
+ "lstrip": false,
56
+ "normalized": false,
57
+ "rstrip": false,
58
+ "single_word": false,
59
+ "special": true
60
+ },
61
+ "151650": {
62
+ "content": "<|quad_start|>",
63
+ "lstrip": false,
64
+ "normalized": false,
65
+ "rstrip": false,
66
+ "single_word": false,
67
+ "special": true
68
+ },
69
+ "151651": {
70
+ "content": "<|quad_end|>",
71
+ "lstrip": false,
72
+ "normalized": false,
73
+ "rstrip": false,
74
+ "single_word": false,
75
+ "special": true
76
+ },
77
+ "151652": {
78
+ "content": "<|vision_start|>",
79
+ "lstrip": false,
80
+ "normalized": false,
81
+ "rstrip": false,
82
+ "single_word": false,
83
+ "special": true
84
+ },
85
+ "151653": {
86
+ "content": "<|vision_end|>",
87
+ "lstrip": false,
88
+ "normalized": false,
89
+ "rstrip": false,
90
+ "single_word": false,
91
+ "special": true
92
+ },
93
+ "151654": {
94
+ "content": "<|vision_pad|>",
95
+ "lstrip": false,
96
+ "normalized": false,
97
+ "rstrip": false,
98
+ "single_word": false,
99
+ "special": true
100
+ },
101
+ "151655": {
102
+ "content": "<|image_pad|>",
103
+ "lstrip": false,
104
+ "normalized": false,
105
+ "rstrip": false,
106
+ "single_word": false,
107
+ "special": true
108
+ },
109
+ "151656": {
110
+ "content": "<|video_pad|>",
111
+ "lstrip": false,
112
+ "normalized": false,
113
+ "rstrip": false,
114
+ "single_word": false,
115
+ "special": true
116
+ },
117
+ "151657": {
118
+ "content": "<tool_call>",
119
+ "lstrip": false,
120
+ "normalized": false,
121
+ "rstrip": false,
122
+ "single_word": false,
123
+ "special": false
124
+ },
125
+ "151658": {
126
+ "content": "</tool_call>",
127
+ "lstrip": false,
128
+ "normalized": false,
129
+ "rstrip": false,
130
+ "single_word": false,
131
+ "special": false
132
+ },
133
+ "151659": {
134
+ "content": "<|fim_prefix|>",
135
+ "lstrip": false,
136
+ "normalized": false,
137
+ "rstrip": false,
138
+ "single_word": false,
139
+ "special": false
140
+ },
141
+ "151660": {
142
+ "content": "<|fim_middle|>",
143
+ "lstrip": false,
144
+ "normalized": false,
145
+ "rstrip": false,
146
+ "single_word": false,
147
+ "special": false
148
+ },
149
+ "151661": {
150
+ "content": "<|fim_suffix|>",
151
+ "lstrip": false,
152
+ "normalized": false,
153
+ "rstrip": false,
154
+ "single_word": false,
155
+ "special": false
156
+ },
157
+ "151662": {
158
+ "content": "<|fim_pad|>",
159
+ "lstrip": false,
160
+ "normalized": false,
161
+ "rstrip": false,
162
+ "single_word": false,
163
+ "special": false
164
+ },
165
+ "151663": {
166
+ "content": "<|repo_name|>",
167
+ "lstrip": false,
168
+ "normalized": false,
169
+ "rstrip": false,
170
+ "single_word": false,
171
+ "special": false
172
+ },
173
+ "151664": {
174
+ "content": "<|file_sep|>",
175
+ "lstrip": false,
176
+ "normalized": false,
177
+ "rstrip": false,
178
+ "single_word": false,
179
+ "special": false
180
+ },
181
+ "151665": {
182
+ "content": "<tool_response>",
183
+ "lstrip": false,
184
+ "normalized": false,
185
+ "rstrip": false,
186
+ "single_word": false,
187
+ "special": false
188
+ },
189
+ "151666": {
190
+ "content": "</tool_response>",
191
+ "lstrip": false,
192
+ "normalized": false,
193
+ "rstrip": false,
194
+ "single_word": false,
195
+ "special": false
196
+ },
197
+ "151667": {
198
+ "content": "<think>",
199
+ "lstrip": false,
200
+ "normalized": false,
201
+ "rstrip": false,
202
+ "single_word": false,
203
+ "special": false
204
+ },
205
+ "151668": {
206
+ "content": "</think>",
207
+ "lstrip": false,
208
+ "normalized": false,
209
+ "rstrip": false,
210
+ "single_word": false,
211
+ "special": false
212
+ }
213
+ },
214
+ "additional_special_tokens": [
215
+ "<|im_start|>",
216
+ "<|im_end|>",
217
+ "<|object_ref_start|>",
218
+ "<|object_ref_end|>",
219
+ "<|box_start|>",
220
+ "<|box_end|>",
221
+ "<|quad_start|>",
222
+ "<|quad_end|>",
223
+ "<|vision_start|>",
224
+ "<|vision_end|>",
225
+ "<|vision_pad|>",
226
+ "<|image_pad|>",
227
+ "<|video_pad|>"
228
+ ],
229
+ "bos_token": null,
230
+ "clean_up_tokenization_spaces": false,
231
+ "eos_token": "<|im_end|>",
232
+ "errors": "replace",
233
+ "extra_special_tokens": {},
234
+ "model_max_length": 262144,
235
+ "pad_token": "<|endoftext|>",
236
+ "processor_class": "Qwen3VLProcessor",
237
+ "split_special_tokens": false,
238
+ "tokenizer_class": "Qwen2Tokenizer",
239
+ "unk_token": null
240
+ }
processor/video_preprocessor_config.json ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "crop_size": null,
3
+ "data_format": "channels_first",
4
+ "default_to_square": true,
5
+ "device": null,
6
+ "do_center_crop": null,
7
+ "do_convert_rgb": true,
8
+ "do_normalize": true,
9
+ "do_rescale": true,
10
+ "do_resize": true,
11
+ "do_sample_frames": true,
12
+ "fps": 2,
13
+ "image_mean": [
14
+ 0.5,
15
+ 0.5,
16
+ 0.5
17
+ ],
18
+ "image_std": [
19
+ 0.5,
20
+ 0.5,
21
+ 0.5
22
+ ],
23
+ "input_data_format": null,
24
+ "max_frames": 768,
25
+ "merge_size": 2,
26
+ "min_frames": 4,
27
+ "num_frames": null,
28
+ "pad_size": null,
29
+ "patch_size": 16,
30
+ "processor_class": "Qwen3VLProcessor",
31
+ "resample": 3,
32
+ "rescale_factor": 0.00392156862745098,
33
+ "return_metadata": false,
34
+ "size": {
35
+ "longest_edge": 25165824,
36
+ "shortest_edge": 4096
37
+ },
38
+ "temporal_patch_size": 2,
39
+ "video_metadata": null,
40
+ "video_processor_type": "Qwen3VLVideoProcessor"
41
+ }
processor/vocab.json ADDED
The diff for this file is too large to render. See raw diff
 
special_tokens_map.json ADDED
@@ -0,0 +1,31 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "additional_special_tokens": [
3
+ "<|im_start|>",
4
+ "<|im_end|>",
5
+ "<|object_ref_start|>",
6
+ "<|object_ref_end|>",
7
+ "<|box_start|>",
8
+ "<|box_end|>",
9
+ "<|quad_start|>",
10
+ "<|quad_end|>",
11
+ "<|vision_start|>",
12
+ "<|vision_end|>",
13
+ "<|vision_pad|>",
14
+ "<|image_pad|>",
15
+ "<|video_pad|>"
16
+ ],
17
+ "eos_token": {
18
+ "content": "<|im_end|>",
19
+ "lstrip": false,
20
+ "normalized": false,
21
+ "rstrip": false,
22
+ "single_word": false
23
+ },
24
+ "pad_token": {
25
+ "content": "<|endoftext|>",
26
+ "lstrip": false,
27
+ "normalized": false,
28
+ "rstrip": false,
29
+ "single_word": false
30
+ }
31
+ }
text_encoder/config.json ADDED
@@ -0,0 +1,65 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Qwen3VLForConditionalGeneration"
4
+ ],
5
+ "dtype": "bfloat16",
6
+ "image_token_id": 151655,
7
+ "model_type": "qwen3_vl",
8
+ "text_config": {
9
+ "attention_bias": false,
10
+ "attention_dropout": 0.0,
11
+ "bos_token_id": 151643,
12
+ "dtype": "bfloat16",
13
+ "eos_token_id": 151645,
14
+ "head_dim": 128,
15
+ "hidden_act": "silu",
16
+ "hidden_size": 4096,
17
+ "initializer_range": 0.02,
18
+ "intermediate_size": 12288,
19
+ "max_position_embeddings": 262144,
20
+ "model_type": "qwen3_vl_text",
21
+ "num_attention_heads": 32,
22
+ "num_hidden_layers": 36,
23
+ "num_key_value_heads": 8,
24
+ "pad_token_id": null,
25
+ "rms_norm_eps": 1e-06,
26
+ "rope_parameters": {
27
+ "mrope_interleaved": true,
28
+ "mrope_section": [
29
+ 24,
30
+ 20,
31
+ 20
32
+ ],
33
+ "rope_theta": 5000000,
34
+ "rope_type": "default"
35
+ },
36
+ "use_cache": true,
37
+ "vocab_size": 151936
38
+ },
39
+ "tie_word_embeddings": false,
40
+ "transformers_version": "5.3.0.dev0",
41
+ "video_token_id": 151656,
42
+ "vision_config": {
43
+ "deepstack_visual_indexes": [
44
+ 8,
45
+ 16,
46
+ 24
47
+ ],
48
+ "depth": 27,
49
+ "dtype": "bfloat16",
50
+ "hidden_act": "gelu_pytorch_tanh",
51
+ "hidden_size": 1152,
52
+ "in_channels": 3,
53
+ "initializer_range": 0.02,
54
+ "intermediate_size": 4304,
55
+ "model_type": "qwen3_vl",
56
+ "num_heads": 16,
57
+ "num_position_embeddings": 2304,
58
+ "out_hidden_size": 4096,
59
+ "patch_size": 16,
60
+ "spatial_merge_size": 2,
61
+ "temporal_patch_size": 2
62
+ },
63
+ "vision_end_token_id": 151653,
64
+ "vision_start_token_id": 151652
65
+ }
text_encoder/generation_config.json ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 151643,
3
+ "do_sample": true,
4
+ "eos_token_id": [
5
+ 151645,
6
+ 151643
7
+ ],
8
+ "pad_token_id": 151643,
9
+ "temperature": 0.7,
10
+ "top_k": 20,
11
+ "top_p": 0.8,
12
+ "transformers_version": "5.3.0.dev0"
13
+ }
text_encoder/model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ba4cadb5ccea2f11dcbb7facce01cb35a76c9a2d5553b34f681b2e16227a1f02
3
+ size 16289680712
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:aeb13307a71acd8fe81861d94ad54ab689df773318809eed3cbe794b4492dae4
3
+ size 11422654
tokenizer_config.json ADDED
@@ -0,0 +1,240 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_bos_token": false,
3
+ "add_prefix_space": false,
4
+ "added_tokens_decoder": {
5
+ "151643": {
6
+ "content": "<|endoftext|>",
7
+ "lstrip": false,
8
+ "normalized": false,
9
+ "rstrip": false,
10
+ "single_word": false,
11
+ "special": true
12
+ },
13
+ "151644": {
14
+ "content": "<|im_start|>",
15
+ "lstrip": false,
16
+ "normalized": false,
17
+ "rstrip": false,
18
+ "single_word": false,
19
+ "special": true
20
+ },
21
+ "151645": {
22
+ "content": "<|im_end|>",
23
+ "lstrip": false,
24
+ "normalized": false,
25
+ "rstrip": false,
26
+ "single_word": false,
27
+ "special": true
28
+ },
29
+ "151646": {
30
+ "content": "<|object_ref_start|>",
31
+ "lstrip": false,
32
+ "normalized": false,
33
+ "rstrip": false,
34
+ "single_word": false,
35
+ "special": true
36
+ },
37
+ "151647": {
38
+ "content": "<|object_ref_end|>",
39
+ "lstrip": false,
40
+ "normalized": false,
41
+ "rstrip": false,
42
+ "single_word": false,
43
+ "special": true
44
+ },
45
+ "151648": {
46
+ "content": "<|box_start|>",
47
+ "lstrip": false,
48
+ "normalized": false,
49
+ "rstrip": false,
50
+ "single_word": false,
51
+ "special": true
52
+ },
53
+ "151649": {
54
+ "content": "<|box_end|>",
55
+ "lstrip": false,
56
+ "normalized": false,
57
+ "rstrip": false,
58
+ "single_word": false,
59
+ "special": true
60
+ },
61
+ "151650": {
62
+ "content": "<|quad_start|>",
63
+ "lstrip": false,
64
+ "normalized": false,
65
+ "rstrip": false,
66
+ "single_word": false,
67
+ "special": true
68
+ },
69
+ "151651": {
70
+ "content": "<|quad_end|>",
71
+ "lstrip": false,
72
+ "normalized": false,
73
+ "rstrip": false,
74
+ "single_word": false,
75
+ "special": true
76
+ },
77
+ "151652": {
78
+ "content": "<|vision_start|>",
79
+ "lstrip": false,
80
+ "normalized": false,
81
+ "rstrip": false,
82
+ "single_word": false,
83
+ "special": true
84
+ },
85
+ "151653": {
86
+ "content": "<|vision_end|>",
87
+ "lstrip": false,
88
+ "normalized": false,
89
+ "rstrip": false,
90
+ "single_word": false,
91
+ "special": true
92
+ },
93
+ "151654": {
94
+ "content": "<|vision_pad|>",
95
+ "lstrip": false,
96
+ "normalized": false,
97
+ "rstrip": false,
98
+ "single_word": false,
99
+ "special": true
100
+ },
101
+ "151655": {
102
+ "content": "<|image_pad|>",
103
+ "lstrip": false,
104
+ "normalized": false,
105
+ "rstrip": false,
106
+ "single_word": false,
107
+ "special": true
108
+ },
109
+ "151656": {
110
+ "content": "<|video_pad|>",
111
+ "lstrip": false,
112
+ "normalized": false,
113
+ "rstrip": false,
114
+ "single_word": false,
115
+ "special": true
116
+ },
117
+ "151657": {
118
+ "content": "<tool_call>",
119
+ "lstrip": false,
120
+ "normalized": false,
121
+ "rstrip": false,
122
+ "single_word": false,
123
+ "special": false
124
+ },
125
+ "151658": {
126
+ "content": "</tool_call>",
127
+ "lstrip": false,
128
+ "normalized": false,
129
+ "rstrip": false,
130
+ "single_word": false,
131
+ "special": false
132
+ },
133
+ "151659": {
134
+ "content": "<|fim_prefix|>",
135
+ "lstrip": false,
136
+ "normalized": false,
137
+ "rstrip": false,
138
+ "single_word": false,
139
+ "special": false
140
+ },
141
+ "151660": {
142
+ "content": "<|fim_middle|>",
143
+ "lstrip": false,
144
+ "normalized": false,
145
+ "rstrip": false,
146
+ "single_word": false,
147
+ "special": false
148
+ },
149
+ "151661": {
150
+ "content": "<|fim_suffix|>",
151
+ "lstrip": false,
152
+ "normalized": false,
153
+ "rstrip": false,
154
+ "single_word": false,
155
+ "special": false
156
+ },
157
+ "151662": {
158
+ "content": "<|fim_pad|>",
159
+ "lstrip": false,
160
+ "normalized": false,
161
+ "rstrip": false,
162
+ "single_word": false,
163
+ "special": false
164
+ },
165
+ "151663": {
166
+ "content": "<|repo_name|>",
167
+ "lstrip": false,
168
+ "normalized": false,
169
+ "rstrip": false,
170
+ "single_word": false,
171
+ "special": false
172
+ },
173
+ "151664": {
174
+ "content": "<|file_sep|>",
175
+ "lstrip": false,
176
+ "normalized": false,
177
+ "rstrip": false,
178
+ "single_word": false,
179
+ "special": false
180
+ },
181
+ "151665": {
182
+ "content": "<tool_response>",
183
+ "lstrip": false,
184
+ "normalized": false,
185
+ "rstrip": false,
186
+ "single_word": false,
187
+ "special": false
188
+ },
189
+ "151666": {
190
+ "content": "</tool_response>",
191
+ "lstrip": false,
192
+ "normalized": false,
193
+ "rstrip": false,
194
+ "single_word": false,
195
+ "special": false
196
+ },
197
+ "151667": {
198
+ "content": "<think>",
199
+ "lstrip": false,
200
+ "normalized": false,
201
+ "rstrip": false,
202
+ "single_word": false,
203
+ "special": false
204
+ },
205
+ "151668": {
206
+ "content": "</think>",
207
+ "lstrip": false,
208
+ "normalized": false,
209
+ "rstrip": false,
210
+ "single_word": false,
211
+ "special": false
212
+ }
213
+ },
214
+ "additional_special_tokens": [
215
+ "<|im_start|>",
216
+ "<|im_end|>",
217
+ "<|object_ref_start|>",
218
+ "<|object_ref_end|>",
219
+ "<|box_start|>",
220
+ "<|box_end|>",
221
+ "<|quad_start|>",
222
+ "<|quad_end|>",
223
+ "<|vision_start|>",
224
+ "<|vision_end|>",
225
+ "<|vision_pad|>",
226
+ "<|image_pad|>",
227
+ "<|video_pad|>"
228
+ ],
229
+ "bos_token": null,
230
+ "clean_up_tokenization_spaces": false,
231
+ "eos_token": "<|im_end|>",
232
+ "errors": "replace",
233
+ "extra_special_tokens": {},
234
+ "model_max_length": 262144,
235
+ "pad_token": "<|endoftext|>",
236
+ "processor_class": "Qwen3VLProcessor",
237
+ "split_special_tokens": false,
238
+ "tokenizer_class": "Qwen2Tokenizer",
239
+ "unk_token": null
240
+ }
video_preprocessor_config.json ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "crop_size": null,
3
+ "data_format": "channels_first",
4
+ "default_to_square": true,
5
+ "device": null,
6
+ "do_center_crop": null,
7
+ "do_convert_rgb": true,
8
+ "do_normalize": true,
9
+ "do_rescale": true,
10
+ "do_resize": true,
11
+ "do_sample_frames": true,
12
+ "fps": 2,
13
+ "image_mean": [
14
+ 0.5,
15
+ 0.5,
16
+ 0.5
17
+ ],
18
+ "image_std": [
19
+ 0.5,
20
+ 0.5,
21
+ 0.5
22
+ ],
23
+ "input_data_format": null,
24
+ "max_frames": 768,
25
+ "merge_size": 2,
26
+ "min_frames": 4,
27
+ "num_frames": null,
28
+ "pad_size": null,
29
+ "patch_size": 16,
30
+ "processor_class": "Qwen3VLProcessor",
31
+ "resample": 3,
32
+ "rescale_factor": 0.00392156862745098,
33
+ "return_metadata": false,
34
+ "size": {
35
+ "longest_edge": 25165824,
36
+ "shortest_edge": 4096
37
+ },
38
+ "temporal_patch_size": 2,
39
+ "video_metadata": null,
40
+ "video_processor_type": "Qwen3VLVideoProcessor"
41
+ }
vocab.json ADDED
The diff for this file is too large to render. See raw diff