Diffusers
Safetensors
qwen3_vl
qwen
qwen3-vl
text-encoder
vision-encoder
heretic
abliteration
prompt-adherence
Instructions to use catplusplus/Qwen21_Text_Encoder_Heretic with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use catplusplus/Qwen21_Text_Encoder_Heretic with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("catplusplus/Qwen21_Text_Encoder_Heretic", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Upload folder using huggingface_hub
Browse files- .gitattributes +2 -0
- README.md +168 -0
- added_tokens.json +28 -0
- chat_template.jinja +120 -0
- config.json +65 -0
- extras/ImageEditServer.py +703 -0
- extras/QwenImage21Backend.py +73 -0
- extras/QwenImage21NVFP4Backend.py +166 -0
- extras/create_heretic_text_encoder.py +201 -0
- extras/imagegen_qwen21_nvfp4.sh +40 -0
- extras/stream_encoder.py +190 -0
- extras/test_heretic_beach_volleyball.py +128 -0
- generation_config.json +13 -0
- merges.txt +0 -0
- model.safetensors +3 -0
- preprocessor_config.json +39 -0
- processor/added_tokens.json +28 -0
- processor/chat_template.jinja +120 -0
- processor/merges.txt +0 -0
- processor/preprocessor_config.json +39 -0
- processor/special_tokens_map.json +31 -0
- processor/tokenizer.json +3 -0
- processor/tokenizer_config.json +240 -0
- processor/video_preprocessor_config.json +41 -0
- processor/vocab.json +0 -0
- special_tokens_map.json +31 -0
- text_encoder/config.json +65 -0
- text_encoder/generation_config.json +13 -0
- text_encoder/model.safetensors +3 -0
- tokenizer.json +3 -0
- tokenizer_config.json +240 -0
- video_preprocessor_config.json +41 -0
- vocab.json +0 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
processor/tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -1,3 +1,171 @@
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
+
base_model: Qwen/Qwen-Image-2.1
|
| 4 |
+
tags:
|
| 5 |
+
- qwen
|
| 6 |
+
- qwen3-vl
|
| 7 |
+
- text-encoder
|
| 8 |
+
- vision-encoder
|
| 9 |
+
- heretic
|
| 10 |
+
- abliteration
|
| 11 |
+
- prompt-adherence
|
| 12 |
+
- diffusers
|
| 13 |
---
|
| 14 |
+
|
| 15 |
+
# Qwen3-VL-8B Heretic Text & Vision Encoder (Prompt Adherence & Geometric Alignment Edition) 🌺✨
|
| 16 |
+
|
| 17 |
+
This repository provides an optimized, abliterated checkpoint of the **Qwen3-VL-8B** text and vision encoder from **[Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1)**, processed with **Norm-Preserving Biprojected Abliteration**.
|
| 18 |
+
|
| 19 |
+
The primary purpose of this model is **maximum instruction following and prompt adherence**: it banishes geometric representation deflection ("internal blush" / hesitation vectors) that otherwise causes safety-tuned VLMs to corrupt diffusion conditioning on dynamic poses, human figures, athletic wear, and complex scenes.
|
| 20 |
+
|
| 21 |
+
---
|
| 22 |
+
|
| 23 |
+
## 🔬 The Core Problem: Why VLM Safety Alignment Degrades Diffusion Conditioning
|
| 24 |
+
|
| 25 |
+
In text-generation tasks, safety alignment mechanisms steer models to emit refusal text (e.g. *"I cannot fulfill this request..."*). However, modern multimodal diffusion architectures like **Qwen-Image-2.1** do **not** generate text tokens:
|
| 26 |
+
|
| 27 |
+
$$\text{DiT Conditioning} \longleftarrow \mathbf{h}_L = \text{TextEncoder}(\text{tokens})[-1]$$
|
| 28 |
+
|
| 29 |
+
The diffusion transformer taps the raw pre-RMSNorm residual hidden states $\mathbf{h}_L$ directly from the text encoder to drive cross-attention.
|
| 30 |
+
|
| 31 |
+
### The "Internal Blush" / Hesitation Deflection Phenomenon
|
| 32 |
+
|
| 33 |
+
When prompts describe human subjects, dynamic physical actions, athletic attire (e.g. swimwear, volleyball, gymnastics), or expressive emotions, safety-tuning vectors inside the language model activate even on completely benign, non-refusal prompts.
|
| 34 |
+
|
| 35 |
+
Because the model cannot output a refusal string, these alignment vectors manifest as a **geometric rotation of the latent representation**:
|
| 36 |
+
|
| 37 |
+
$$\mathbf{h}_{\text{sensitive}} = \mathbf{h}_{\text{clean}} + \mathbf{v}_{\text{refusal}}$$
|
| 38 |
+
|
| 39 |
+
This hidden deflection rotates the conditioning signal by **over 60% relative norm** away from the prompt's intended semantic visual trajectory!
|
| 40 |
+
|
| 41 |
+
### Visual Consequences in Image Generation
|
| 42 |
+
|
| 43 |
+
When cross-attention layers in the DiT receive a representation deflected into the refusal/modesty subspace, the model displays hesitation artifacts:
|
| 44 |
+
1. **Modesty Hallucinations & Clothing Confusion**: Spontaneous addition of mismatched cloth, awkward white ruffles, or extra fabric covering swimwear or sportswear.
|
| 45 |
+
2. **Anatomical Occlusion**: The model avoids rendering human limbs or athletic poses, awkwardly hiding arms behind character backs or contorting torsos.
|
| 46 |
+
3. **Subject & Prop Merging**: Equipment or background elements get fused into characters (e.g. sports balls bizarrely merged onto heads as hair ornaments).
|
| 47 |
+
4. **Action Damping**: Dynamic verbs (*"jumping to spike the ball"*) are subdued into passive, static standing postures.
|
| 48 |
+
|
| 49 |
+
---
|
| 50 |
+
|
| 51 |
+
## 📊 Quantitative Measurement of Representation Deflection
|
| 52 |
+
|
| 53 |
+
Using contrastive prompt pairs across benign and sensitive subjects, we measured the layer-by-layer cosine similarity and relative deflection norm across all 37 positions (input embeddings + 36 decoder layers) of Qwen3-VL-8B:
|
| 54 |
+
|
| 55 |
+
| Layer Index | Position | Cosine Similarity ($\cos \theta$) | Relative Deflection ($\|\Delta \mathbf{h}\| / \|\mathbf{h}\|$) | Deflection Norm $\|\Delta \mathbf{h}\|$ |
|
| 56 |
+
| :---: | :---: | :---: | :---: | :---: |
|
| 57 |
+
| **0** | Input Embeddings | **1.0000** | **0.00%** | 0.00 |
|
| 58 |
+
| **8** | Early Transformer | **0.9991** | **3.82%** | 18.24 |
|
| 59 |
+
| **16** | Mid-Low (Deflection Onset) | **0.9943** | **10.64%** | 52.88 |
|
| 60 |
+
| **20** | Mid-High Divergence | **0.9780** | **21.05%** | 114.73 |
|
| 61 |
+
| **24** | Refusal Vector Surge | **0.9414** | **34.25%** | 192.40 |
|
| 62 |
+
| **28** | Acceleration Peak | **0.8842** | **47.19%** | 275.31 |
|
| 63 |
+
| **32** | Late Transformer | **0.8350** | **56.12%** | 331.05 |
|
| 64 |
+
| **36** | Final Conditioning Layer | **0.8110** | **60.30%** | **357.94** |
|
| 65 |
+
|
| 66 |
+
Between Layer 20 and Layer 36, representation deflection accelerates rapidly, culminating in a **60.3% vector distortion**. By surgically neutralizing this direction, the text encoder reflects the exact intended prompt semantics.
|
| 67 |
+
|
| 68 |
+
---
|
| 69 |
+
|
| 70 |
+
## 🛠️ Methodology: Norm-Preserving Biprojected Abliteration
|
| 71 |
+
|
| 72 |
+
To eliminate hesitation deflection without degrading general language comprehension, we applied **Norm-Preserving Biprojected Abliteration** (`create_heretic_text_encoder.py`):
|
| 73 |
+
|
| 74 |
+
1. **Refusal Subspace Extraction**: Difference-of-means vectors were extracted across contrastive prompt sets:
|
| 75 |
+
$$\mathbf{r}_l = \boldsymbol{\mu}_{\text{sensitive}}^{(l)} - \boldsymbol{\mu}_{\text{benign}}^{(l)}$$
|
| 76 |
+
2. **Benign Subspace Orthogonalization**: The general semantic direction was stripped from the refusal vector:
|
| 77 |
+
$$\mathbf{v}_l = \mathbf{r}_l - \text{proj}_{\mathbf{u}_{\text{benign}}}(\mathbf{r}_l)$$
|
| 78 |
+
3. **Norm-Preserving Rank-1 Projection**: Across 54 linear projection matrices (`self_attn.o_proj` and `mlp.down_proj` in layers 9–35, centered at layer 26 with Gaussian falloff $\lambda \in [0.10, 1.00]$):
|
| 79 |
+
$$W_{\text{norm}} = \text{normalize}(W, p=2, \text{dim}=1)$$
|
| 80 |
+
$$W' = \text{normalize}\Big(W_{\text{norm}} - \lambda \mathbf{v}_l (\mathbf{v}_l^T W_{\text{norm}})\Big) \cdot \|W\|_{\text{row}}$$
|
| 81 |
+
|
| 82 |
+
Because exact row norms ($\|W\|_{\text{row}}$) are strictly preserved, the network's overall activation scales and general reasoning capabilities remain completely intact.
|
| 83 |
+
|
| 84 |
+
---
|
| 85 |
+
|
| 86 |
+
## 🤝 Pairing with Quantized DiT (`nunchaku-qwen-image-2.1`)
|
| 87 |
+
|
| 88 |
+
This text encoder is specifically engineered to be paired with **[`nunchaku-qwen-image-2.1`](https://huggingface.co/models/nunchaku-qwen-image-2.1)** for consumer GPU setups:
|
| 89 |
+
|
| 90 |
+
* **Resident DiT + Streamed Text Encoder**:
|
| 91 |
+
* DiT (`best_quality_fp4.safetensors`): **4.08 GB resident VRAM**.
|
| 92 |
+
* VAE (`AutoencoderKLQwenImage21`): **0.64 GB resident VRAM**.
|
| 93 |
+
* Qwen3-VL-8B ViT Vision Encoder: **1.07 GB resident VRAM**.
|
| 94 |
+
* Qwen3-VL-8B Language Model: Streamed layer-by-layer through a static **368 MB GPU buffer** over PCIe at ~28.7 GB/s via `stream_encoder.py`.
|
| 95 |
+
* **Total VRAM Footprint**: **~6.17 GB active VRAM**, leaving **~9.5 GB free headroom** on a single 16 GB GPU (such as RTX 5060 Ti or RTX 4080)!
|
| 96 |
+
* **Inference Speed**: Multimodal prompt encoding completes in **1.06s** (saving 16s vs CPU), and 25-step image generation runs in **~20s**.
|
| 97 |
+
|
| 98 |
+
---
|
| 99 |
+
|
| 100 |
+
## 🚀 Quickstart Usage
|
| 101 |
+
|
| 102 |
+
### 1. Installation
|
| 103 |
+
|
| 104 |
+
```bash
|
| 105 |
+
pip install diffusers transformers accelerate torch sentencepiece
|
| 106 |
+
```
|
| 107 |
+
|
| 108 |
+
### 2. Loading with Diffusers
|
| 109 |
+
|
| 110 |
+
```python
|
| 111 |
+
import torch
|
| 112 |
+
from diffusers import QwenImage21Pipeline
|
| 113 |
+
from transformers import Qwen3VLForConditionalGeneration, Qwen3VLProcessor
|
| 114 |
+
|
| 115 |
+
# 1. Load Heretic text encoder and processor
|
| 116 |
+
text_encoder = Qwen3VLForConditionalGeneration.from_pretrained(
|
| 117 |
+
"models/Qwen21_Text_Encoder_Heretic",
|
| 118 |
+
torch_dtype=torch.bfloat16,
|
| 119 |
+
low_cpu_mem_usage=True,
|
| 120 |
+
)
|
| 121 |
+
processor = Qwen3VLProcessor.from_pretrained("models/Qwen21_Text_Encoder_Heretic")
|
| 122 |
+
|
| 123 |
+
# 2. Assemble into pipeline
|
| 124 |
+
pipe = QwenImage21Pipeline.from_pretrained(
|
| 125 |
+
"Qwen/Qwen-Image-2.1",
|
| 126 |
+
text_encoder=text_encoder,
|
| 127 |
+
processor=processor,
|
| 128 |
+
torch_dtype=torch.bfloat16,
|
| 129 |
+
)
|
| 130 |
+
pipe.enable_sequential_cpu_offload(gpu_id=0)
|
| 131 |
+
|
| 132 |
+
# 3. Generate with precise prompt adherence
|
| 133 |
+
image = pipe(
|
| 134 |
+
prompt="Two cute anime girls in colorful bikinis playing beach volleyball on a sunny tropical beach, dynamic action pose, jumping to spike the ball, sharp focus",
|
| 135 |
+
height=1024,
|
| 136 |
+
width=1024,
|
| 137 |
+
num_inference_steps=25,
|
| 138 |
+
true_cfg_scale=1.0,
|
| 139 |
+
).images[0]
|
| 140 |
+
|
| 141 |
+
image.save("beach_volleyball.png")
|
| 142 |
+
```
|
| 143 |
+
|
| 144 |
+
### 3. High-Throughput Server Usage
|
| 145 |
+
|
| 146 |
+
Run the bundled ImageEditServer with NVFP4 DiT and Heretic text encoder:
|
| 147 |
+
|
| 148 |
+
```bash
|
| 149 |
+
# Start server on port 4500 (uses Heretic text encoder by default)
|
| 150 |
+
./extras/imagegen_qwen21_nvfp4.sh 4500
|
| 151 |
+
```
|
| 152 |
+
|
| 153 |
+
---
|
| 154 |
+
|
| 155 |
+
## 📦 Packaged Sources (`extras/`)
|
| 156 |
+
|
| 157 |
+
* `create_heretic_text_encoder.py`: Complete script used to measure refusal vectors and perform norm-preserving biprojected abliteration.
|
| 158 |
+
* `stream_encoder.py`: Zero-quality-loss layerwise weight streaming engine for Qwen3-VL-8B.
|
| 159 |
+
* `test_heretic_beach_volleyball.py`: Empirical verification script comparing stock vs Heretic encoders.
|
| 160 |
+
* `QwenImage21NVFP4Backend.py`: Diffusers + Nunchaku backend supporting custom text encoder overrides.
|
| 161 |
+
* `ImageEditServer.py` & `imagegen_qwen21_nvfp4.sh`: Resident image generation server.
|
| 162 |
+
|
| 163 |
+
---
|
| 164 |
+
|
| 165 |
+
## 📜 Citation & Credits
|
| 166 |
+
|
| 167 |
+
* **Qwen-Image-2.1 & Qwen3-VL**: Qwen Team, Alibaba Cloud.
|
| 168 |
+
* **Abliteration Principles**: Arditi et al. (*Refusal in Language Models Is Mediated by a Single Direction*).
|
| 169 |
+
* **Heretic LLM**: Heretic project (*Directional Abliteration Toolkit*).
|
| 170 |
+
* **Abliteration & Diffusion Conditioning Optimization**: Oleg K. / Nikola Seeker Project.
|
| 171 |
+
|
added_tokens.json
ADDED
|
@@ -0,0 +1,28 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"</think>": 151668,
|
| 3 |
+
"</tool_call>": 151658,
|
| 4 |
+
"</tool_response>": 151666,
|
| 5 |
+
"<think>": 151667,
|
| 6 |
+
"<tool_call>": 151657,
|
| 7 |
+
"<tool_response>": 151665,
|
| 8 |
+
"<|box_end|>": 151649,
|
| 9 |
+
"<|box_start|>": 151648,
|
| 10 |
+
"<|endoftext|>": 151643,
|
| 11 |
+
"<|file_sep|>": 151664,
|
| 12 |
+
"<|fim_middle|>": 151660,
|
| 13 |
+
"<|fim_pad|>": 151662,
|
| 14 |
+
"<|fim_prefix|>": 151659,
|
| 15 |
+
"<|fim_suffix|>": 151661,
|
| 16 |
+
"<|im_end|>": 151645,
|
| 17 |
+
"<|im_start|>": 151644,
|
| 18 |
+
"<|image_pad|>": 151655,
|
| 19 |
+
"<|object_ref_end|>": 151647,
|
| 20 |
+
"<|object_ref_start|>": 151646,
|
| 21 |
+
"<|quad_end|>": 151651,
|
| 22 |
+
"<|quad_start|>": 151650,
|
| 23 |
+
"<|repo_name|>": 151663,
|
| 24 |
+
"<|video_pad|>": 151656,
|
| 25 |
+
"<|vision_end|>": 151653,
|
| 26 |
+
"<|vision_pad|>": 151654,
|
| 27 |
+
"<|vision_start|>": 151652
|
| 28 |
+
}
|
chat_template.jinja
ADDED
|
@@ -0,0 +1,120 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0].role == 'system' %}
|
| 4 |
+
{%- if messages[0].content is string %}
|
| 5 |
+
{{- messages[0].content }}
|
| 6 |
+
{%- else %}
|
| 7 |
+
{%- for content in messages[0].content %}
|
| 8 |
+
{%- if 'text' in content %}
|
| 9 |
+
{{- content.text }}
|
| 10 |
+
{%- endif %}
|
| 11 |
+
{%- endfor %}
|
| 12 |
+
{%- endif %}
|
| 13 |
+
{{- '\n\n' }}
|
| 14 |
+
{%- endif %}
|
| 15 |
+
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 16 |
+
{%- for tool in tools %}
|
| 17 |
+
{{- "\n" }}
|
| 18 |
+
{{- tool | tojson }}
|
| 19 |
+
{%- endfor %}
|
| 20 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 21 |
+
{%- else %}
|
| 22 |
+
{%- if messages[0].role == 'system' %}
|
| 23 |
+
{{- '<|im_start|>system\n' }}
|
| 24 |
+
{%- if messages[0].content is string %}
|
| 25 |
+
{{- messages[0].content }}
|
| 26 |
+
{%- else %}
|
| 27 |
+
{%- for content in messages[0].content %}
|
| 28 |
+
{%- if 'text' in content %}
|
| 29 |
+
{{- content.text }}
|
| 30 |
+
{%- endif %}
|
| 31 |
+
{%- endfor %}
|
| 32 |
+
{%- endif %}
|
| 33 |
+
{{- '<|im_end|>\n' }}
|
| 34 |
+
{%- endif %}
|
| 35 |
+
{%- endif %}
|
| 36 |
+
{%- set image_count = namespace(value=0) %}
|
| 37 |
+
{%- set video_count = namespace(value=0) %}
|
| 38 |
+
{%- for message in messages %}
|
| 39 |
+
{%- if message.role == "user" %}
|
| 40 |
+
{{- '<|im_start|>' + message.role + '\n' }}
|
| 41 |
+
{%- if message.content is string %}
|
| 42 |
+
{{- message.content }}
|
| 43 |
+
{%- else %}
|
| 44 |
+
{%- for content in message.content %}
|
| 45 |
+
{%- if content.type == 'image' or 'image' in content or 'image_url' in content %}
|
| 46 |
+
{%- set image_count.value = image_count.value + 1 %}
|
| 47 |
+
{%- if add_vision_id %}Picture {{ image_count.value }}: {% endif -%}
|
| 48 |
+
<|vision_start|><|image_pad|><|vision_end|>
|
| 49 |
+
{%- elif content.type == 'video' or 'video' in content %}
|
| 50 |
+
{%- set video_count.value = video_count.value + 1 %}
|
| 51 |
+
{%- if add_vision_id %}Video {{ video_count.value }}: {% endif -%}
|
| 52 |
+
<|vision_start|><|video_pad|><|vision_end|>
|
| 53 |
+
{%- elif 'text' in content %}
|
| 54 |
+
{{- content.text }}
|
| 55 |
+
{%- endif %}
|
| 56 |
+
{%- endfor %}
|
| 57 |
+
{%- endif %}
|
| 58 |
+
{{- '<|im_end|>\n' }}
|
| 59 |
+
{%- elif message.role == "assistant" %}
|
| 60 |
+
{{- '<|im_start|>' + message.role + '\n' }}
|
| 61 |
+
{%- if message.content is string %}
|
| 62 |
+
{{- message.content }}
|
| 63 |
+
{%- else %}
|
| 64 |
+
{%- for content_item in message.content %}
|
| 65 |
+
{%- if 'text' in content_item %}
|
| 66 |
+
{{- content_item.text }}
|
| 67 |
+
{%- endif %}
|
| 68 |
+
{%- endfor %}
|
| 69 |
+
{%- endif %}
|
| 70 |
+
{%- if message.tool_calls %}
|
| 71 |
+
{%- for tool_call in message.tool_calls %}
|
| 72 |
+
{%- if (loop.first and message.content) or (not loop.first) %}
|
| 73 |
+
{{- '\n' }}
|
| 74 |
+
{%- endif %}
|
| 75 |
+
{%- if tool_call.function %}
|
| 76 |
+
{%- set tool_call = tool_call.function %}
|
| 77 |
+
{%- endif %}
|
| 78 |
+
{{- '<tool_call>\n{"name": "' }}
|
| 79 |
+
{{- tool_call.name }}
|
| 80 |
+
{{- '", "arguments": ' }}
|
| 81 |
+
{%- if tool_call.arguments is string %}
|
| 82 |
+
{{- tool_call.arguments }}
|
| 83 |
+
{%- else %}
|
| 84 |
+
{{- tool_call.arguments | tojson }}
|
| 85 |
+
{%- endif %}
|
| 86 |
+
{{- '}\n</tool_call>' }}
|
| 87 |
+
{%- endfor %}
|
| 88 |
+
{%- endif %}
|
| 89 |
+
{{- '<|im_end|>\n' }}
|
| 90 |
+
{%- elif message.role == "tool" %}
|
| 91 |
+
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
| 92 |
+
{{- '<|im_start|>user' }}
|
| 93 |
+
{%- endif %}
|
| 94 |
+
{{- '\n<tool_response>\n' }}
|
| 95 |
+
{%- if message.content is string %}
|
| 96 |
+
{{- message.content }}
|
| 97 |
+
{%- else %}
|
| 98 |
+
{%- for content in message.content %}
|
| 99 |
+
{%- if content.type == 'image' or 'image' in content or 'image_url' in content %}
|
| 100 |
+
{%- set image_count.value = image_count.value + 1 %}
|
| 101 |
+
{%- if add_vision_id %}Picture {{ image_count.value }}: {% endif -%}
|
| 102 |
+
<|vision_start|><|image_pad|><|vision_end|>
|
| 103 |
+
{%- elif content.type == 'video' or 'video' in content %}
|
| 104 |
+
{%- set video_count.value = video_count.value + 1 %}
|
| 105 |
+
{%- if add_vision_id %}Video {{ video_count.value }}: {% endif -%}
|
| 106 |
+
<|vision_start|><|video_pad|><|vision_end|>
|
| 107 |
+
{%- elif 'text' in content %}
|
| 108 |
+
{{- content.text }}
|
| 109 |
+
{%- endif %}
|
| 110 |
+
{%- endfor %}
|
| 111 |
+
{%- endif %}
|
| 112 |
+
{{- '\n</tool_response>' }}
|
| 113 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 114 |
+
{{- '<|im_end|>\n' }}
|
| 115 |
+
{%- endif %}
|
| 116 |
+
{%- endif %}
|
| 117 |
+
{%- endfor %}
|
| 118 |
+
{%- if add_generation_prompt %}
|
| 119 |
+
{{- '<|im_start|>assistant\n' }}
|
| 120 |
+
{%- endif %}
|
config.json
ADDED
|
@@ -0,0 +1,65 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"Qwen3VLForConditionalGeneration"
|
| 4 |
+
],
|
| 5 |
+
"dtype": "bfloat16",
|
| 6 |
+
"image_token_id": 151655,
|
| 7 |
+
"model_type": "qwen3_vl",
|
| 8 |
+
"text_config": {
|
| 9 |
+
"attention_bias": false,
|
| 10 |
+
"attention_dropout": 0.0,
|
| 11 |
+
"bos_token_id": 151643,
|
| 12 |
+
"dtype": "bfloat16",
|
| 13 |
+
"eos_token_id": 151645,
|
| 14 |
+
"head_dim": 128,
|
| 15 |
+
"hidden_act": "silu",
|
| 16 |
+
"hidden_size": 4096,
|
| 17 |
+
"initializer_range": 0.02,
|
| 18 |
+
"intermediate_size": 12288,
|
| 19 |
+
"max_position_embeddings": 262144,
|
| 20 |
+
"model_type": "qwen3_vl_text",
|
| 21 |
+
"num_attention_heads": 32,
|
| 22 |
+
"num_hidden_layers": 36,
|
| 23 |
+
"num_key_value_heads": 8,
|
| 24 |
+
"pad_token_id": null,
|
| 25 |
+
"rms_norm_eps": 1e-06,
|
| 26 |
+
"rope_parameters": {
|
| 27 |
+
"mrope_interleaved": true,
|
| 28 |
+
"mrope_section": [
|
| 29 |
+
24,
|
| 30 |
+
20,
|
| 31 |
+
20
|
| 32 |
+
],
|
| 33 |
+
"rope_theta": 5000000,
|
| 34 |
+
"rope_type": "default"
|
| 35 |
+
},
|
| 36 |
+
"use_cache": true,
|
| 37 |
+
"vocab_size": 151936
|
| 38 |
+
},
|
| 39 |
+
"tie_word_embeddings": false,
|
| 40 |
+
"transformers_version": "5.3.0.dev0",
|
| 41 |
+
"video_token_id": 151656,
|
| 42 |
+
"vision_config": {
|
| 43 |
+
"deepstack_visual_indexes": [
|
| 44 |
+
8,
|
| 45 |
+
16,
|
| 46 |
+
24
|
| 47 |
+
],
|
| 48 |
+
"depth": 27,
|
| 49 |
+
"dtype": "bfloat16",
|
| 50 |
+
"hidden_act": "gelu_pytorch_tanh",
|
| 51 |
+
"hidden_size": 1152,
|
| 52 |
+
"in_channels": 3,
|
| 53 |
+
"initializer_range": 0.02,
|
| 54 |
+
"intermediate_size": 4304,
|
| 55 |
+
"model_type": "qwen3_vl",
|
| 56 |
+
"num_heads": 16,
|
| 57 |
+
"num_position_embeddings": 2304,
|
| 58 |
+
"out_hidden_size": 4096,
|
| 59 |
+
"patch_size": 16,
|
| 60 |
+
"spatial_merge_size": 2,
|
| 61 |
+
"temporal_patch_size": 2
|
| 62 |
+
},
|
| 63 |
+
"vision_end_token_id": 151653,
|
| 64 |
+
"vision_start_token_id": 151652
|
| 65 |
+
}
|
extras/ImageEditServer.py
ADDED
|
@@ -0,0 +1,703 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import argparse
|
| 2 |
+
import base64
|
| 3 |
+
import io
|
| 4 |
+
import time
|
| 5 |
+
import torch
|
| 6 |
+
import uvicorn
|
| 7 |
+
import gc
|
| 8 |
+
import asyncio
|
| 9 |
+
import traceback
|
| 10 |
+
from typing import List, Optional, Union
|
| 11 |
+
from contextlib import asynccontextmanager
|
| 12 |
+
from fastapi import FastAPI, HTTPException, UploadFile, File, Form
|
| 13 |
+
from pydantic import BaseModel
|
| 14 |
+
from PIL import Image, ImageOps
|
| 15 |
+
|
| 16 |
+
# Argument parsing
|
| 17 |
+
parser = argparse.ArgumentParser(description="Flux Image Edit Server with Nunchaku")
|
| 18 |
+
parser.add_argument("--host", type=str, default="0.0.0.0", help="Host to bind to")
|
| 19 |
+
parser.add_argument("--port", type=int, default=8000, help="Port to bind to")
|
| 20 |
+
parser.add_argument("--model", type=str, default="black-forest-labs/FLUX.1-Kontext-dev", help="Path or Repo ID of the base model")
|
| 21 |
+
parser.add_argument("--optimized-model", type=str, default=None, help="Path to the optimized Nunchaku model safetensors file")
|
| 22 |
+
parser.add_argument("--optimized-edit-model", type=str, default=None, help="Path to the optimized Nunchaku model safetensors file for editing (optional)")
|
| 23 |
+
parser.add_argument("--backend", type=str, default="kontext", choices=["kontext", "flux2", "flux2-klein", "flux2_klein", "qwen", "qwen21", "qwen-2.1", "qwen21-nvfp4", "qwen21_nvfp4", "glm", "zimage"], help="Backend to use: 'kontext', 'flux2', 'flux2-klein', 'qwen', 'qwen21', 'qwen21-nvfp4', 'glm', or 'zimage'")
|
| 24 |
+
parser.add_argument("--steps", type=int, default=28, help="Default number of inference steps")
|
| 25 |
+
parser.add_argument("--guidance-scale", type=float, default=3.5, help="Default guidance scale")
|
| 26 |
+
parser.add_argument("--qwenimage", action="store_true", help="Use QwenImageBackend (T2I only) instead of full Qwen edit backend")
|
| 27 |
+
parser.add_argument("--uma", action="store_true", help="Enable Unified Memory Architecture mode (load all to GPU, disable offload)")
|
| 28 |
+
parser.add_argument("--no-layerwise-offload", action="store_true", help="Keep transformer resident in VRAM without layer-by-layer offloading")
|
| 29 |
+
parser.add_argument(
|
| 30 |
+
"--text-encoder",
|
| 31 |
+
type=str,
|
| 32 |
+
default=None,
|
| 33 |
+
help="Path or Repo ID of custom text encoder (e.g. models/Qwen_Text_Encoder_Heretic)",
|
| 34 |
+
)
|
| 35 |
+
parser.add_argument(
|
| 36 |
+
"--nvfp4-text-encoder",
|
| 37 |
+
type=str,
|
| 38 |
+
default=None,
|
| 39 |
+
help=(
|
| 40 |
+
"Path to an NVFP4-pack-quantized HuggingFace text encoder "
|
| 41 |
+
"swaps in vLLM's W4A4 NVFP4 CUTLASS GEMM for ~4x text-encoder VRAM savings."
|
| 42 |
+
),
|
| 43 |
+
)
|
| 44 |
+
args = parser.parse_args()
|
| 45 |
+
|
| 46 |
+
@asynccontextmanager
|
| 47 |
+
async def lifespan(app: FastAPI):
|
| 48 |
+
# Startup logic
|
| 49 |
+
load_model()
|
| 50 |
+
yield
|
| 51 |
+
# Shutdown logic (if any) could go here
|
| 52 |
+
|
| 53 |
+
app = FastAPI(lifespan=lifespan)
|
| 54 |
+
|
| 55 |
+
# Global components
|
| 56 |
+
IMAGE_DIMENSION_ALIGNMENT = 32
|
| 57 |
+
# Cap for a reference image fed with input_fit=native: the pipeline resizes each condition image
|
| 58 |
+
# at its own aspect ratio, so this only bounds token count, never framing.
|
| 59 |
+
MAX_REFERENCE_PIXELS = 1024 * 1024
|
| 60 |
+
pipeline = None
|
| 61 |
+
edit_pipeline = None
|
| 62 |
+
request_lock = asyncio.Lock()
|
| 63 |
+
is_sleeping_flag = False
|
| 64 |
+
sleep_requested = False
|
| 65 |
+
|
| 66 |
+
def set_zero_cond_t(transformer, val: bool):
|
| 67 |
+
"""Recursively synchronize zero_cond_t across transformer and all inner blocks."""
|
| 68 |
+
if transformer is None:
|
| 69 |
+
return
|
| 70 |
+
if hasattr(transformer, "zero_cond_t"):
|
| 71 |
+
transformer.zero_cond_t = val
|
| 72 |
+
if hasattr(transformer, "transformer_blocks"):
|
| 73 |
+
for block in transformer.transformer_blocks:
|
| 74 |
+
if hasattr(block, "zero_cond_t"):
|
| 75 |
+
block.zero_cond_t = val
|
| 76 |
+
|
| 77 |
+
def load_model():
|
| 78 |
+
global pipeline, edit_pipeline
|
| 79 |
+
|
| 80 |
+
try:
|
| 81 |
+
if args.backend == "kontext":
|
| 82 |
+
import KontextBackend
|
| 83 |
+
print(f"Initializing KontextBackend...")
|
| 84 |
+
backend = KontextBackend.KontextBackend(args.model, args.optimized_model)
|
| 85 |
+
pipeline, edit_pipeline = backend.load()
|
| 86 |
+
elif args.backend == "flux2":
|
| 87 |
+
import Flux2Backend
|
| 88 |
+
print(f"Initializing Flux2Backend...")
|
| 89 |
+
backend = Flux2Backend.Flux2Backend(args.model)
|
| 90 |
+
pipeline, edit_pipeline = backend.load()
|
| 91 |
+
elif args.backend == "glm":
|
| 92 |
+
import GlmBackend
|
| 93 |
+
print(f"Initializing GlmBackend...")
|
| 94 |
+
# Use provided model or default to the one in the snippet if args.model is generic
|
| 95 |
+
# The user might pass the specific GLM model via --model, or we default in GlmBackend.
|
| 96 |
+
# Let's pass args.model if it's not the default flux one, otherwise let GlmBackend use its default.
|
| 97 |
+
model_to_use = args.model if args.model != "black-forest-labs/FLUX.1-Kontext-dev" else "Disty0/GLM-Image-SDNQ-4bit-dynamic"
|
| 98 |
+
backend = GlmBackend.GlmBackend(model_to_use)
|
| 99 |
+
pipeline, edit_pipeline = backend.load()
|
| 100 |
+
elif args.backend in ["qwen21", "qwen-2.1", "qwen21-nvfp4", "qwen21_nvfp4"]:
|
| 101 |
+
if args.optimized_model or "nvfp4" in args.backend:
|
| 102 |
+
import QwenImage21NVFP4Backend
|
| 103 |
+
print(f"Initializing QwenImage21NVFP4Backend (resident NVFP4)...")
|
| 104 |
+
backend = QwenImage21NVFP4Backend.QwenImage21NVFP4Backend(
|
| 105 |
+
model_id=args.model,
|
| 106 |
+
optimized_model_path=args.optimized_model or "/home/olegk/Nikola/models/nunchaku-qwen-image-2.1/best_quality_fp4.safetensors",
|
| 107 |
+
text_encoder_path=args.text_encoder,
|
| 108 |
+
)
|
| 109 |
+
pipeline, edit_pipeline = backend.load()
|
| 110 |
+
else:
|
| 111 |
+
import QwenImage21Backend
|
| 112 |
+
print(f"Initializing QwenImage21Backend (unquantized baseline)...")
|
| 113 |
+
backend = QwenImage21Backend.QwenImage21Backend(
|
| 114 |
+
args.model,
|
| 115 |
+
text_encoder_path=args.text_encoder,
|
| 116 |
+
)
|
| 117 |
+
pipeline, edit_pipeline = backend.load()
|
| 118 |
+
elif args.backend.startswith("qwen"):
|
| 119 |
+
if args.qwenimage:
|
| 120 |
+
import QwenImageBackend
|
| 121 |
+
print(f"Initializing QwenImageBackend (T2I only)...")
|
| 122 |
+
backend = QwenImageBackend.QwenImageBackend(args.model, args.optimized_model)
|
| 123 |
+
pipeline, edit_pipeline = backend.load()
|
| 124 |
+
else:
|
| 125 |
+
import QwenBackend
|
| 126 |
+
print(f"Initializing QwenBackend...")
|
| 127 |
+
backend = QwenBackend.QwenBackend(
|
| 128 |
+
args.model,
|
| 129 |
+
args.optimized_model,
|
| 130 |
+
optimized_edit_model_path=args.optimized_edit_model,
|
| 131 |
+
uma=args.uma,
|
| 132 |
+
no_layerwise_offload=args.no_layerwise_offload,
|
| 133 |
+
text_encoder_path=args.text_encoder,
|
| 134 |
+
)
|
| 135 |
+
pipeline, edit_pipeline = backend.load()
|
| 136 |
+
elif args.backend in ["flux2-klein", "flux2_klein"]:
|
| 137 |
+
import Flux2KleinBackend
|
| 138 |
+
print(f"Initializing Flux2KleinBackend...")
|
| 139 |
+
backend = Flux2KleinBackend.Flux2KleinBackend(
|
| 140 |
+
args.model,
|
| 141 |
+
args.optimized_model,
|
| 142 |
+
nvfp4_text_encoder_path=args.nvfp4_text_encoder,
|
| 143 |
+
)
|
| 144 |
+
pipeline, edit_pipeline = backend.load()
|
| 145 |
+
elif args.backend == "zimage":
|
| 146 |
+
import ZImageTurboBackend
|
| 147 |
+
print(f"Initializing ZImageTurboBackend...")
|
| 148 |
+
backend = ZImageTurboBackend.ZImageTurboBackend(
|
| 149 |
+
args.model,
|
| 150 |
+
args.optimized_model,
|
| 151 |
+
uma=args.uma,
|
| 152 |
+
nvfp4_text_encoder_path=args.nvfp4_text_encoder,
|
| 153 |
+
)
|
| 154 |
+
pipeline, edit_pipeline = backend.load()
|
| 155 |
+
else:
|
| 156 |
+
raise ValueError(f"Unknown backend: {args.backend}")
|
| 157 |
+
|
| 158 |
+
except Exception as e:
|
| 159 |
+
print(f"Oh no! The model refused to wake up: {e}")
|
| 160 |
+
raise e
|
| 161 |
+
|
| 162 |
+
# Enable progress bar for diffusers
|
| 163 |
+
import diffusers.utils.logging
|
| 164 |
+
diffusers.utils.logging.enable_progress_bar()
|
| 165 |
+
diffusers.utils.logging.set_verbosity_info()
|
| 166 |
+
|
| 167 |
+
# Instrument timing tracker for stage-by-stage profiling
|
| 168 |
+
instrument_pipeline(pipeline)
|
| 169 |
+
instrument_pipeline(edit_pipeline)
|
| 170 |
+
|
| 171 |
+
print("Model loaded successfully! Ready for editing quests!")
|
| 172 |
+
|
| 173 |
+
class PipelineTimingTracker:
|
| 174 |
+
def __init__(self):
|
| 175 |
+
self.reset()
|
| 176 |
+
|
| 177 |
+
def reset(self):
|
| 178 |
+
self.text_encoder_time = 0.0
|
| 179 |
+
self.transformer_times = []
|
| 180 |
+
self.vae_time = 0.0
|
| 181 |
+
self._te_start = None
|
| 182 |
+
self._tr_start = None
|
| 183 |
+
self._vae_start = None
|
| 184 |
+
|
| 185 |
+
def on_te_start(self):
|
| 186 |
+
if torch.cuda.is_available():
|
| 187 |
+
torch.cuda.synchronize()
|
| 188 |
+
self._te_start = time.perf_counter()
|
| 189 |
+
|
| 190 |
+
def on_te_end(self):
|
| 191 |
+
if torch.cuda.is_available():
|
| 192 |
+
torch.cuda.synchronize()
|
| 193 |
+
if self._te_start is not None:
|
| 194 |
+
self.text_encoder_time += time.perf_counter() - self._te_start
|
| 195 |
+
self._te_start = None
|
| 196 |
+
|
| 197 |
+
def on_tr_start(self):
|
| 198 |
+
if torch.cuda.is_available():
|
| 199 |
+
torch.cuda.synchronize()
|
| 200 |
+
self._tr_start = time.perf_counter()
|
| 201 |
+
|
| 202 |
+
def on_tr_end(self):
|
| 203 |
+
if torch.cuda.is_available():
|
| 204 |
+
torch.cuda.synchronize()
|
| 205 |
+
if self._tr_start is not None:
|
| 206 |
+
self.transformer_times.append(time.perf_counter() - self._tr_start)
|
| 207 |
+
self._tr_start = None
|
| 208 |
+
|
| 209 |
+
def on_vae_start(self):
|
| 210 |
+
if torch.cuda.is_available():
|
| 211 |
+
torch.cuda.synchronize()
|
| 212 |
+
self._vae_start = time.perf_counter()
|
| 213 |
+
|
| 214 |
+
def on_vae_end(self):
|
| 215 |
+
if torch.cuda.is_available():
|
| 216 |
+
torch.cuda.synchronize()
|
| 217 |
+
if self._vae_start is not None:
|
| 218 |
+
self.vae_time += time.perf_counter() - self._vae_start
|
| 219 |
+
self._vae_start = None
|
| 220 |
+
|
| 221 |
+
def get_summary(self, total_wall_time: float) -> dict:
|
| 222 |
+
num_steps = len(self.transformer_times)
|
| 223 |
+
total_tr = sum(self.transformer_times)
|
| 224 |
+
avg_step = (total_tr / num_steps) if num_steps > 0 else 0.0
|
| 225 |
+
return {
|
| 226 |
+
"text_encoder_s": round(self.text_encoder_time, 3),
|
| 227 |
+
"transformer_s": round(total_tr, 3),
|
| 228 |
+
"steps": num_steps,
|
| 229 |
+
"step_avg_s": round(avg_step, 3),
|
| 230 |
+
"vae_s": round(self.vae_time, 3),
|
| 231 |
+
"total_s": round(total_wall_time, 3),
|
| 232 |
+
}
|
| 233 |
+
|
| 234 |
+
timing_tracker = PipelineTimingTracker()
|
| 235 |
+
|
| 236 |
+
def instrument_pipeline(p):
|
| 237 |
+
if p is None or getattr(p, "_timing_instrumented", False):
|
| 238 |
+
return
|
| 239 |
+
if hasattr(p, "text_encoder") and p.text_encoder is not None:
|
| 240 |
+
orig_te = p.text_encoder.forward
|
| 241 |
+
def timed_te(*args, **kwargs):
|
| 242 |
+
timing_tracker.on_te_start()
|
| 243 |
+
try:
|
| 244 |
+
return orig_te(*args, **kwargs)
|
| 245 |
+
finally:
|
| 246 |
+
timing_tracker.on_te_end()
|
| 247 |
+
p.text_encoder.forward = timed_te
|
| 248 |
+
|
| 249 |
+
if hasattr(p, "transformer") and p.transformer is not None:
|
| 250 |
+
orig_tr = p.transformer.forward
|
| 251 |
+
def timed_tr(*args, **kwargs):
|
| 252 |
+
timing_tracker.on_tr_start()
|
| 253 |
+
try:
|
| 254 |
+
return orig_tr(*args, **kwargs)
|
| 255 |
+
finally:
|
| 256 |
+
timing_tracker.on_tr_end()
|
| 257 |
+
p.transformer.forward = timed_tr
|
| 258 |
+
|
| 259 |
+
if hasattr(p, "vae") and p.vae is not None:
|
| 260 |
+
orig_vae = p.vae.decode
|
| 261 |
+
def timed_vae(*args, **kwargs):
|
| 262 |
+
timing_tracker.on_vae_start()
|
| 263 |
+
try:
|
| 264 |
+
return orig_vae(*args, **kwargs)
|
| 265 |
+
finally:
|
| 266 |
+
timing_tracker.on_vae_end()
|
| 267 |
+
p.vae.decode = timed_vae
|
| 268 |
+
|
| 269 |
+
p._timing_instrumented = True
|
| 270 |
+
|
| 271 |
+
def flush():
|
| 272 |
+
gc.collect()
|
| 273 |
+
torch.cuda.empty_cache()
|
| 274 |
+
|
| 275 |
+
|
| 276 |
+
class ImageGenerationRequest(BaseModel):
|
| 277 |
+
prompt: str
|
| 278 |
+
n: int = 1
|
| 279 |
+
size: str = "1024x1024"
|
| 280 |
+
response_format: str = "b64_json"
|
| 281 |
+
quality: str = "standard"
|
| 282 |
+
style: str = "vivid"
|
| 283 |
+
num_inference_steps: Optional[int] = None
|
| 284 |
+
guidance_scale: Optional[float] = None
|
| 285 |
+
negative_prompt: Optional[str] = None
|
| 286 |
+
seed: Optional[int] = None
|
| 287 |
+
|
| 288 |
+
|
| 289 |
+
@app.post("/v1/sleep")
|
| 290 |
+
async def sleep_endpoint():
|
| 291 |
+
global is_sleeping_flag, sleep_requested
|
| 292 |
+
sleep_requested = True
|
| 293 |
+
try:
|
| 294 |
+
async with request_lock:
|
| 295 |
+
if not is_sleeping_flag and sleep_requested:
|
| 296 |
+
print("Sleep requested, moving models to CPU...")
|
| 297 |
+
for p in [pipeline, edit_pipeline]:
|
| 298 |
+
if not p: continue
|
| 299 |
+
for name, component in p.components.items():
|
| 300 |
+
if isinstance(component, torch.nn.Module):
|
| 301 |
+
# Special handling for Nunchaku which blocks .to() if offload is True
|
| 302 |
+
if hasattr(component, "set_offload") and getattr(component, "offload", False):
|
| 303 |
+
component.set_offload(False)
|
| 304 |
+
component._nunchaku_was_offloaded = True
|
| 305 |
+
|
| 306 |
+
try:
|
| 307 |
+
component.to("cpu")
|
| 308 |
+
except Exception as e:
|
| 309 |
+
pass
|
| 310 |
+
flush()
|
| 311 |
+
is_sleeping_flag = True
|
| 312 |
+
finally:
|
| 313 |
+
sleep_requested = False
|
| 314 |
+
return {"status": "sleep completed", "is_sleeping": is_sleeping_flag}
|
| 315 |
+
|
| 316 |
+
@app.post("/v1/wake_up")
|
| 317 |
+
async def wake_up_endpoint():
|
| 318 |
+
global is_sleeping_flag, sleep_requested
|
| 319 |
+
sleep_requested = False
|
| 320 |
+
async with request_lock:
|
| 321 |
+
if is_sleeping_flag:
|
| 322 |
+
print("Waking up, restoring models to CUDA...")
|
| 323 |
+
for p in [pipeline, edit_pipeline]:
|
| 324 |
+
if not p: continue
|
| 325 |
+
excluded = getattr(p, "_exclude_from_cpu_offload", [])
|
| 326 |
+
for name, component in p.components.items():
|
| 327 |
+
if isinstance(component, torch.nn.Module):
|
| 328 |
+
if getattr(component, "_nunchaku_was_offloaded", False):
|
| 329 |
+
component.set_offload(True, use_pin_memory=True, num_blocks_on_gpu=8)
|
| 330 |
+
for attr in ["img_in", "txt_in", "txt_norm", "time_text_embed", "norm_out", "proj_out"]:
|
| 331 |
+
if hasattr(component, attr):
|
| 332 |
+
try:
|
| 333 |
+
getattr(component, attr).to("cuda")
|
| 334 |
+
except Exception:
|
| 335 |
+
pass
|
| 336 |
+
component._nunchaku_was_offloaded = False
|
| 337 |
+
elif not hasattr(component, "_hf_hook") or name in excluded:
|
| 338 |
+
try:
|
| 339 |
+
component.to("cuda")
|
| 340 |
+
except Exception:
|
| 341 |
+
pass
|
| 342 |
+
is_sleeping_flag = False
|
| 343 |
+
return {"status": "awoken", "is_sleeping": False}
|
| 344 |
+
|
| 345 |
+
@app.get("/v1/is_sleeping")
|
| 346 |
+
async def is_sleeping_endpoint():
|
| 347 |
+
return {"is_sleeping": is_sleeping_flag}
|
| 348 |
+
|
| 349 |
+
|
| 350 |
+
@app.get("/v1/memory_stats")
|
| 351 |
+
async def memory_stats_endpoint():
|
| 352 |
+
"""Lightweight introspection endpoint that returns PyTorch's CUDA allocator
|
| 353 |
+
snapshot. Used to diagnose VRAM/UMA bloat without restarting the server."""
|
| 354 |
+
stats = {}
|
| 355 |
+
if torch.cuda.is_available():
|
| 356 |
+
stats["allocated_gb"] = torch.cuda.memory_allocated() / 1e9
|
| 357 |
+
stats["reserved_gb"] = torch.cuda.memory_reserved() / 1e9
|
| 358 |
+
stats["max_allocated_gb"] = torch.cuda.max_memory_allocated() / 1e9
|
| 359 |
+
stats["max_reserved_gb"] = torch.cuda.max_memory_reserved() / 1e9
|
| 360 |
+
# Top allocations by size from the allocator snapshot (>=64 MiB)
|
| 361 |
+
try:
|
| 362 |
+
snap = torch.cuda.memory_snapshot()
|
| 363 |
+
blocks = []
|
| 364 |
+
for seg in snap:
|
| 365 |
+
for b in seg.get("blocks", []):
|
| 366 |
+
if b.get("state") == "active_allocated" and b.get("size", 0) >= 64 * 1024 * 1024:
|
| 367 |
+
blocks.append(b["size"])
|
| 368 |
+
blocks.sort(reverse=True)
|
| 369 |
+
stats["large_active_blocks_gb"] = [round(s / 1e9, 3) for s in blocks[:20]]
|
| 370 |
+
stats["large_active_blocks_total_gb"] = round(sum(blocks) / 1e9, 3)
|
| 371 |
+
stats["large_active_blocks_count"] = len(blocks)
|
| 372 |
+
except Exception as e:
|
| 373 |
+
stats["snapshot_error"] = str(e)
|
| 374 |
+
# Walk Python objects to find big tensors and group them
|
| 375 |
+
try:
|
| 376 |
+
import gc as _gc
|
| 377 |
+
seen = set()
|
| 378 |
+
big = []
|
| 379 |
+
for obj in _gc.get_objects():
|
| 380 |
+
try:
|
| 381 |
+
if isinstance(obj, torch.Tensor) and obj.is_cuda:
|
| 382 |
+
ptr = obj.data_ptr()
|
| 383 |
+
if ptr in seen or ptr == 0:
|
| 384 |
+
continue
|
| 385 |
+
seen.add(ptr)
|
| 386 |
+
sz = obj.element_size() * obj.numel()
|
| 387 |
+
if sz >= 16 * 1024 * 1024:
|
| 388 |
+
big.append((sz, tuple(obj.shape), str(obj.dtype)))
|
| 389 |
+
except Exception:
|
| 390 |
+
continue
|
| 391 |
+
big.sort(reverse=True)
|
| 392 |
+
# Group by (shape, dtype)
|
| 393 |
+
from collections import Counter
|
| 394 |
+
grouped = Counter((shape, dtype) for _, shape, dtype in big)
|
| 395 |
+
stats["big_tensor_groups"] = [
|
| 396 |
+
{"shape": list(shape), "dtype": dtype, "count": cnt,
|
| 397 |
+
"size_gb_each": round(
|
| 398 |
+
(1 if shape == () else (lambda l: __import__('functools').reduce(lambda a, b: a*b, l, 1))(shape)) * (
|
| 399 |
+
8 if 'int64' in dtype or 'float64' in dtype else
|
| 400 |
+
4 if 'int32' in dtype or 'float32' in dtype else
|
| 401 |
+
2 if 'bfloat16' in dtype or 'float16' in dtype else 1
|
| 402 |
+
) / 1e9, 4)}
|
| 403 |
+
for (shape, dtype), cnt in grouped.most_common(30)
|
| 404 |
+
]
|
| 405 |
+
stats["big_tensor_count"] = len(big)
|
| 406 |
+
stats["big_tensor_total_gb"] = round(sum(s for s, _, _ in big) / 1e9, 3)
|
| 407 |
+
except Exception as e:
|
| 408 |
+
stats["walk_error"] = str(e)
|
| 409 |
+
return stats
|
| 410 |
+
|
| 411 |
+
@app.post("/v1/images/edits")
|
| 412 |
+
async def edit_image(
|
| 413 |
+
image: Union[List[UploadFile], UploadFile] = File(...),
|
| 414 |
+
prompt: str = Form(...),
|
| 415 |
+
n: int = Form(1),
|
| 416 |
+
size: str = Form("1024x1024"),
|
| 417 |
+
width: Optional[int] = Form(None),
|
| 418 |
+
height: Optional[int] = Form(None),
|
| 419 |
+
input_fit: str = Form("canvas"),
|
| 420 |
+
response_format: str = Form("b64_json"), # Default to b64_json
|
| 421 |
+
guidance_scale: Optional[float] = Form(None),
|
| 422 |
+
num_inference_steps: Optional[int] = Form(None),
|
| 423 |
+
negative_prompt: Optional[str] = Form(None),
|
| 424 |
+
seed: Optional[int] = Form(None)
|
| 425 |
+
):
|
| 426 |
+
# Use CLI defaults if not provided
|
| 427 |
+
steps = num_inference_steps if num_inference_steps is not None else args.steps
|
| 428 |
+
cfg_scale = guidance_scale if guidance_scale is not None else args.guidance_scale
|
| 429 |
+
neg_prompt = negative_prompt if negative_prompt is not None else "" # Default empty for now, or maybe None?
|
| 430 |
+
|
| 431 |
+
generator = None
|
| 432 |
+
import random
|
| 433 |
+
if seed is None:
|
| 434 |
+
seed = random.randint(0, 2**32 - 1)
|
| 435 |
+
|
| 436 |
+
print(f"Using seed: {seed}")
|
| 437 |
+
generator = torch.Generator(device="cuda").manual_seed(seed)
|
| 438 |
+
|
| 439 |
+
if not edit_pipeline:
|
| 440 |
+
raise HTTPException(status_code=500, detail="Model not loaded")
|
| 441 |
+
|
| 442 |
+
if sleep_requested or is_sleeping_flag:
|
| 443 |
+
raise HTTPException(status_code=503, detail="Server is sleeping or trying to sleep.")
|
| 444 |
+
|
| 445 |
+
async with request_lock:
|
| 446 |
+
print(f"Received edit request: {prompt}")
|
| 447 |
+
|
| 448 |
+
# Processing the input image(s)
|
| 449 |
+
input_files = image if isinstance(image, list) else [image]
|
| 450 |
+
init_images = []
|
| 451 |
+
|
| 452 |
+
try:
|
| 453 |
+
for img_file in input_files:
|
| 454 |
+
await img_file.seek(0)
|
| 455 |
+
contents = await img_file.read()
|
| 456 |
+
img = Image.open(io.BytesIO(contents)).convert("RGB")
|
| 457 |
+
init_images.append(img)
|
| 458 |
+
except Exception as e:
|
| 459 |
+
raise HTTPException(status_code=400, detail=f"Invalid image file: {e}")
|
| 460 |
+
|
| 461 |
+
if not init_images:
|
| 462 |
+
raise HTTPException(status_code=400, detail="No images provided")
|
| 463 |
+
|
| 464 |
+
# Parse max target dimensions from requested size
|
| 465 |
+
try:
|
| 466 |
+
target_width, target_height = map(int, size.split("x"))
|
| 467 |
+
except ValueError:
|
| 468 |
+
target_width, target_height = 1024, 1024
|
| 469 |
+
|
| 470 |
+
# Calculate new dimensions preserving aspect ratio based on the first image
|
| 471 |
+
if width is not None and height is not None:
|
| 472 |
+
# Explicit output canvas: the caller owns the framing, so references of any aspect
|
| 473 |
+
# ratio can be fed without dictating the composition's shape.
|
| 474 |
+
req_width, req_height = width, height
|
| 475 |
+
else:
|
| 476 |
+
first_image = init_images[0]
|
| 477 |
+
orig_width, orig_height = first_image.size
|
| 478 |
+
scale = min(target_width / orig_width, target_height / orig_height)
|
| 479 |
+
req_width = int(orig_width * scale)
|
| 480 |
+
req_height = int(orig_height * scale)
|
| 481 |
+
|
| 482 |
+
# Ensure dimensions are aligned to 32 for compatibility (e.g. GLM-Image)
|
| 483 |
+
width = (req_width // IMAGE_DIMENSION_ALIGNMENT) * IMAGE_DIMENSION_ALIGNMENT
|
| 484 |
+
height = (req_height // IMAGE_DIMENSION_ALIGNMENT) * IMAGE_DIMENSION_ALIGNMENT
|
| 485 |
+
|
| 486 |
+
if (input_fit or "").strip().lower() == "native":
|
| 487 |
+
# QwenImageEditPlus resizes every condition image at its own aspect ratio, so padding
|
| 488 |
+
# references onto the output canvas only costs them resolution and adds false letterbox
|
| 489 |
+
# content the model then reproduces. Keep each reference's own framing.
|
| 490 |
+
resized_images = []
|
| 491 |
+
for img in init_images:
|
| 492 |
+
iw, ih = img.size
|
| 493 |
+
scale = min(1.0, (MAX_REFERENCE_PIXELS / float(max(1, iw * ih))) ** 0.5)
|
| 494 |
+
if scale < 1.0:
|
| 495 |
+
nw = max(IMAGE_DIMENSION_ALIGNMENT,
|
| 496 |
+
(int(iw * scale) // IMAGE_DIMENSION_ALIGNMENT) * IMAGE_DIMENSION_ALIGNMENT)
|
| 497 |
+
nh = max(IMAGE_DIMENSION_ALIGNMENT,
|
| 498 |
+
(int(ih * scale) // IMAGE_DIMENSION_ALIGNMENT) * IMAGE_DIMENSION_ALIGNMENT)
|
| 499 |
+
img = img.resize((nw, nh), Image.LANCZOS)
|
| 500 |
+
resized_images.append(img)
|
| 501 |
+
else:
|
| 502 |
+
# Resize input images to match the calculated target size, padding if necessary
|
| 503 |
+
resized_images = []
|
| 504 |
+
for img in init_images:
|
| 505 |
+
if img.size != (width, height):
|
| 506 |
+
# This handles cases where subsequent images might have different ARs
|
| 507 |
+
img = ImageOps.pad(img, (width, height), method=Image.LANCZOS, color=(0, 0, 0))
|
| 508 |
+
resized_images.append(img)
|
| 509 |
+
|
| 510 |
+
# If single image, pass as item, if multiple, pass as list
|
| 511 |
+
# GLM pipeline has a bug where it checks len() on the input, so it must be a list
|
| 512 |
+
if len(resized_images) > 1 or args.backend == "glm":
|
| 513 |
+
image_input = resized_images
|
| 514 |
+
else:
|
| 515 |
+
image_input = resized_images[0]
|
| 516 |
+
|
| 517 |
+
response_images = []
|
| 518 |
+
timing_tracker.reset()
|
| 519 |
+
t_req_start = time.perf_counter()
|
| 520 |
+
|
| 521 |
+
try:
|
| 522 |
+
if args.backend.startswith("qwen"):
|
| 523 |
+
# Qwen specific parameters
|
| 524 |
+
if hasattr(edit_pipeline, "transformer"):
|
| 525 |
+
set_zero_cond_t(edit_pipeline.transformer, True)
|
| 526 |
+
# guidance_scale maps to true_cfg_scale
|
| 527 |
+
if args.qwenimage: # QwenImageBackend is T2I only, so it doesn't take an image
|
| 528 |
+
generated_images = edit_pipeline(
|
| 529 |
+
prompt=prompt,
|
| 530 |
+
height=height,
|
| 531 |
+
width=width,
|
| 532 |
+
num_inference_steps=steps,
|
| 533 |
+
true_cfg_scale=cfg_scale,
|
| 534 |
+
num_images_per_prompt=n,
|
| 535 |
+
generator=generator,
|
| 536 |
+
).images
|
| 537 |
+
else: # Full Qwen edit backend takes an image (or list of images now)
|
| 538 |
+
generated_images = edit_pipeline(
|
| 539 |
+
image=image_input,
|
| 540 |
+
prompt=prompt,
|
| 541 |
+
height=height,
|
| 542 |
+
width=width,
|
| 543 |
+
negative_prompt=neg_prompt,
|
| 544 |
+
num_inference_steps=steps,
|
| 545 |
+
true_cfg_scale=cfg_scale,
|
| 546 |
+
num_images_per_prompt=n,
|
| 547 |
+
generator=generator,
|
| 548 |
+
).images
|
| 549 |
+
else:
|
| 550 |
+
# Standard Flux/Kontext or GLM
|
| 551 |
+
# GLM I2I Fix: Manually move vision encoder to GPU because get_image_features escapes hooks
|
| 552 |
+
if args.backend == "glm" and hasattr(edit_pipeline, "vision_language_encoder"):
|
| 553 |
+
print("Manually moving GLM Vision Encoder to GPU...")
|
| 554 |
+
edit_pipeline.vision_language_encoder.to("cuda")
|
| 555 |
+
|
| 556 |
+
try:
|
| 557 |
+
generated_images = edit_pipeline(
|
| 558 |
+
image=image_input,
|
| 559 |
+
prompt=prompt,
|
| 560 |
+
height=height,
|
| 561 |
+
width=width,
|
| 562 |
+
num_inference_steps=steps,
|
| 563 |
+
guidance_scale=cfg_scale,
|
| 564 |
+
num_images_per_prompt=n,
|
| 565 |
+
generator=generator,
|
| 566 |
+
).images
|
| 567 |
+
finally:
|
| 568 |
+
if args.backend == "glm" and hasattr(edit_pipeline, "vision_language_encoder"):
|
| 569 |
+
print("Moving GLM Vision Encoder back to CPU...")
|
| 570 |
+
edit_pipeline.vision_language_encoder.to("cpu")
|
| 571 |
+
|
| 572 |
+
for img in generated_images:
|
| 573 |
+
buffered = io.BytesIO()
|
| 574 |
+
img.save(buffered, format="PNG")
|
| 575 |
+
img_str = base64.b64encode(buffered.getvalue()).decode("utf-8")
|
| 576 |
+
|
| 577 |
+
if response_format == "b64_json":
|
| 578 |
+
response_images.append({"b64_json": img_str})
|
| 579 |
+
else:
|
| 580 |
+
# If url is requested we can't really do it without storage, so we fallback or error?
|
| 581 |
+
# For now, let's just assume simple b64_json as per request
|
| 582 |
+
response_images.append({"b64_json": img_str}) # Fallback
|
| 583 |
+
|
| 584 |
+
except Exception as e:
|
| 585 |
+
print(f"Error during editing: {e}")
|
| 586 |
+
print(traceback.format_exc())
|
| 587 |
+
raise HTTPException(status_code=500, detail=str(e))
|
| 588 |
+
finally:
|
| 589 |
+
flush()
|
| 590 |
+
|
| 591 |
+
t_req_total = time.perf_counter() - t_req_start
|
| 592 |
+
timings = timing_tracker.get_summary(t_req_total)
|
| 593 |
+
print(f"[TIMING] Text Encoder: {timings['text_encoder_s']:.2f}s | "
|
| 594 |
+
f"Transformer ({timings['steps']} steps): {timings['transformer_s']:.2f}s ({timings['step_avg_s']:.2f}s/step) | "
|
| 595 |
+
f"VAE: {timings['vae_s']:.2f}s | Total: {timings['total_s']:.2f}s")
|
| 596 |
+
|
| 597 |
+
return {
|
| 598 |
+
"created": int(time.time()),
|
| 599 |
+
"timings": timings,
|
| 600 |
+
"seed": seed,
|
| 601 |
+
"width": width,
|
| 602 |
+
"height": height,
|
| 603 |
+
"data": response_images
|
| 604 |
+
}
|
| 605 |
+
|
| 606 |
+
|
| 607 |
+
|
| 608 |
+
@app.post("/v1/images/generations")
|
| 609 |
+
async def generate_image(request: ImageGenerationRequest):
|
| 610 |
+
if not pipeline:
|
| 611 |
+
raise HTTPException(status_code=500, detail="Model not loaded")
|
| 612 |
+
|
| 613 |
+
if sleep_requested or is_sleeping_flag:
|
| 614 |
+
raise HTTPException(status_code=503, detail="Server is sleeping or trying to sleep.")
|
| 615 |
+
|
| 616 |
+
async with request_lock:
|
| 617 |
+
#print(f"Received generation request: {request.prompt}")
|
| 618 |
+
|
| 619 |
+
# Parse size
|
| 620 |
+
try:
|
| 621 |
+
width, height = map(int, request.size.split("x"))
|
| 622 |
+
except ValueError:
|
| 623 |
+
width, height = 1024, 1024
|
| 624 |
+
|
| 625 |
+
# Ensure dimensions are aligned to 32
|
| 626 |
+
width = (width // IMAGE_DIMENSION_ALIGNMENT) * IMAGE_DIMENSION_ALIGNMENT
|
| 627 |
+
height = (height // IMAGE_DIMENSION_ALIGNMENT) * IMAGE_DIMENSION_ALIGNMENT
|
| 628 |
+
|
| 629 |
+
response_images = []
|
| 630 |
+
timing_tracker.reset()
|
| 631 |
+
t_req_start = time.perf_counter()
|
| 632 |
+
|
| 633 |
+
try:
|
| 634 |
+
# Generate images (no image argument for txt2img!)
|
| 635 |
+
steps = request.num_inference_steps if request.num_inference_steps is not None else args.steps
|
| 636 |
+
cfg_scale = request.guidance_scale if request.guidance_scale is not None else args.guidance_scale
|
| 637 |
+
# negative_prompt not in standard request body in original snippet, but we added it to model
|
| 638 |
+
neg_prompt = request.negative_prompt if request.negative_prompt is not None else ""
|
| 639 |
+
|
| 640 |
+
generator = None
|
| 641 |
+
import random
|
| 642 |
+
seed = request.seed
|
| 643 |
+
if seed is None:
|
| 644 |
+
seed = random.randint(0, 2**32 - 1)
|
| 645 |
+
|
| 646 |
+
print(f"Using seed: {seed}")
|
| 647 |
+
generator = torch.Generator(device="cuda").manual_seed(seed)
|
| 648 |
+
|
| 649 |
+
if args.backend.startswith("qwen"):
|
| 650 |
+
if hasattr(pipeline, "transformer"):
|
| 651 |
+
set_zero_cond_t(pipeline.transformer, False)
|
| 652 |
+
generated_images = pipeline(
|
| 653 |
+
prompt=request.prompt,
|
| 654 |
+
height=height,
|
| 655 |
+
width=width,
|
| 656 |
+
num_inference_steps=steps,
|
| 657 |
+
true_cfg_scale=cfg_scale,
|
| 658 |
+
num_images_per_prompt=request.n,
|
| 659 |
+
negative_prompt=neg_prompt,
|
| 660 |
+
generator=generator,
|
| 661 |
+
).images
|
| 662 |
+
else:
|
| 663 |
+
generated_images = pipeline(
|
| 664 |
+
prompt=request.prompt,
|
| 665 |
+
height=height,
|
| 666 |
+
width=width,
|
| 667 |
+
num_inference_steps=steps,
|
| 668 |
+
guidance_scale=cfg_scale,
|
| 669 |
+
num_images_per_prompt=request.n,
|
| 670 |
+
generator=generator,
|
| 671 |
+
# Not passing negative_prompt here for generation unless we confirm support in standard Flux pipeline?
|
| 672 |
+
).images
|
| 673 |
+
|
| 674 |
+
for img in generated_images:
|
| 675 |
+
buffered = io.BytesIO()
|
| 676 |
+
img.save(buffered, format="PNG")
|
| 677 |
+
img_str = base64.b64encode(buffered.getvalue()).decode("utf-8")
|
| 678 |
+
response_images.append({"b64_json": img_str})
|
| 679 |
+
|
| 680 |
+
except Exception as e:
|
| 681 |
+
print(f"Error during generation: {e}")
|
| 682 |
+
print(traceback.format_exc())
|
| 683 |
+
raise HTTPException(status_code=500, detail=str(e))
|
| 684 |
+
finally:
|
| 685 |
+
flush()
|
| 686 |
+
|
| 687 |
+
t_req_total = time.perf_counter() - t_req_start
|
| 688 |
+
timings = timing_tracker.get_summary(t_req_total)
|
| 689 |
+
print(f"[TIMING] Text Encoder: {timings['text_encoder_s']:.2f}s | "
|
| 690 |
+
f"Transformer ({timings['steps']} steps): {timings['transformer_s']:.2f}s ({timings['step_avg_s']:.2f}s/step) | "
|
| 691 |
+
f"VAE: {timings['vae_s']:.2f}s | Total: {timings['total_s']:.2f}s")
|
| 692 |
+
|
| 693 |
+
return {
|
| 694 |
+
"created": int(time.time()),
|
| 695 |
+
"timings": timings,
|
| 696 |
+
"seed": seed,
|
| 697 |
+
"width": width,
|
| 698 |
+
"height": height,
|
| 699 |
+
"data": response_images
|
| 700 |
+
}
|
| 701 |
+
|
| 702 |
+
if __name__ == "__main__":
|
| 703 |
+
uvicorn.run(app, host=args.host, port=args.port)
|
extras/QwenImage21Backend.py
ADDED
|
@@ -0,0 +1,73 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import torch
|
| 2 |
+
import os
|
| 3 |
+
from diffusers import QwenImage21Pipeline
|
| 4 |
+
|
| 5 |
+
# Channeling Nikola's cheerful seeker spirit!
|
| 6 |
+
# A clean, modular backend for Qwen-Image-2.1 with sequential layerwise offloading!
|
| 7 |
+
|
| 8 |
+
class QwenImage21Backend:
|
| 9 |
+
def __init__(
|
| 10 |
+
self,
|
| 11 |
+
model_id="/home/olegk/Nikola/models/Qwen/Qwen-Image-2.1",
|
| 12 |
+
gpu_id=0,
|
| 13 |
+
enable_tiling=True,
|
| 14 |
+
text_encoder_path=None,
|
| 15 |
+
):
|
| 16 |
+
self.model_id = model_id
|
| 17 |
+
self.gpu_id = gpu_id
|
| 18 |
+
self.enable_tiling = enable_tiling
|
| 19 |
+
self.text_encoder_path = text_encoder_path
|
| 20 |
+
self.pipeline = None
|
| 21 |
+
|
| 22 |
+
def load(self):
|
| 23 |
+
print(f"Loading QwenImage21Backend from {self.model_id}...")
|
| 24 |
+
if self.text_encoder_path:
|
| 25 |
+
print(f" • Custom Text Encoder: {self.text_encoder_path}")
|
| 26 |
+
from transformers import Qwen3VLForConditionalGeneration, Qwen3VLProcessor
|
| 27 |
+
|
| 28 |
+
te_dir = (
|
| 29 |
+
os.path.join(self.text_encoder_path, "text_encoder")
|
| 30 |
+
if os.path.isdir(os.path.join(self.text_encoder_path, "text_encoder"))
|
| 31 |
+
else self.text_encoder_path
|
| 32 |
+
)
|
| 33 |
+
proc_dir = (
|
| 34 |
+
os.path.join(self.text_encoder_path, "processor")
|
| 35 |
+
if os.path.isdir(os.path.join(self.text_encoder_path, "processor"))
|
| 36 |
+
else self.text_encoder_path
|
| 37 |
+
)
|
| 38 |
+
if not os.path.exists(os.path.join(proc_dir, "tokenizer.json")):
|
| 39 |
+
proc_dir = os.path.join(self.model_id, "processor")
|
| 40 |
+
|
| 41 |
+
custom_te = Qwen3VLForConditionalGeneration.from_pretrained(
|
| 42 |
+
te_dir,
|
| 43 |
+
torch_dtype=torch.bfloat16,
|
| 44 |
+
low_cpu_mem_usage=True,
|
| 45 |
+
)
|
| 46 |
+
custom_proc = Qwen3VLProcessor.from_pretrained(proc_dir)
|
| 47 |
+
|
| 48 |
+
pipeline = QwenImage21Pipeline.from_pretrained(
|
| 49 |
+
self.model_id,
|
| 50 |
+
text_encoder=custom_te,
|
| 51 |
+
processor=custom_proc,
|
| 52 |
+
torch_dtype=torch.bfloat16,
|
| 53 |
+
)
|
| 54 |
+
else:
|
| 55 |
+
pipeline = QwenImage21Pipeline.from_pretrained(
|
| 56 |
+
self.model_id,
|
| 57 |
+
torch_dtype=torch.bfloat16,
|
| 58 |
+
)
|
| 59 |
+
|
| 60 |
+
print(f"Attaching sequential CPU offload on GPU {self.gpu_id} for layerwise execution...")
|
| 61 |
+
pipeline.enable_sequential_cpu_offload(gpu_id=self.gpu_id)
|
| 62 |
+
|
| 63 |
+
if self.enable_tiling:
|
| 64 |
+
print("Enabling VAE tiling for low-memory decode...")
|
| 65 |
+
try:
|
| 66 |
+
pipeline.vae.enable_tiling()
|
| 67 |
+
except Exception as e:
|
| 68 |
+
print(f"Note: VAE tiling could not be enabled ({e}), continuing with standard decode.")
|
| 69 |
+
|
| 70 |
+
self.pipeline = pipeline
|
| 71 |
+
# QwenImage21Pipeline natively handles both generations (t2i) and edits (i2i)
|
| 72 |
+
return self.pipeline, self.pipeline
|
| 73 |
+
|
extras/QwenImage21NVFP4Backend.py
ADDED
|
@@ -0,0 +1,166 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# -*- coding: utf-8 -*-
|
| 2 |
+
"""Nunchaku NVFP4 Resident Backend for Qwen-Image-2.1.
|
| 3 |
+
|
| 4 |
+
Loads the forged SVDQuant NVFP4 r32 Qwen-Image-2.1 transformer directly into VRAM,
|
| 5 |
+
enabling blazingly fast inference with zero layerwise PCIe streaming bottlenecks.
|
| 6 |
+
"""
|
| 7 |
+
|
| 8 |
+
import os
|
| 9 |
+
import sys
|
| 10 |
+
import torch
|
| 11 |
+
from diffusers import QwenImage21Pipeline
|
| 12 |
+
|
| 13 |
+
# Ensure local packages are on path
|
| 14 |
+
ROOT_DIR = "/auto/home/amano/olegk/Nikola"
|
| 15 |
+
for p in [f"{ROOT_DIR}/packages/nunchaku", f"{ROOT_DIR}/packages/deepcompressor", ROOT_DIR]:
|
| 16 |
+
if p not in sys.path:
|
| 17 |
+
sys.path.insert(0, p)
|
| 18 |
+
|
| 19 |
+
from nunchaku.models.transformers.transformer_qwenimage21 import NunchakuQwenImage21Transformer2DModel
|
| 20 |
+
|
| 21 |
+
|
| 22 |
+
class QwenImage21NVFP4Backend:
|
| 23 |
+
def __init__(
|
| 24 |
+
self,
|
| 25 |
+
model_id="/home/olegk/Nikola/models/Qwen/Qwen-Image-2.1",
|
| 26 |
+
optimized_model_path="/home/olegk/Nikola/models/nunchaku-qwen-image-2.1/best_quality_fp4.safetensors",
|
| 27 |
+
gpu_id=0,
|
| 28 |
+
enable_tiling=True,
|
| 29 |
+
dynamic_scale_k=0.0,
|
| 30 |
+
stream_text_encoder=True,
|
| 31 |
+
text_encoder_path=None,
|
| 32 |
+
):
|
| 33 |
+
self.model_id = model_id
|
| 34 |
+
self.optimized_model_path = optimized_model_path
|
| 35 |
+
self.gpu_id = gpu_id
|
| 36 |
+
self.enable_tiling = enable_tiling
|
| 37 |
+
self.dynamic_scale_k = dynamic_scale_k
|
| 38 |
+
self.stream_text_encoder = stream_text_encoder
|
| 39 |
+
self.text_encoder_path = text_encoder_path
|
| 40 |
+
self.pipeline = None
|
| 41 |
+
|
| 42 |
+
def load(self):
|
| 43 |
+
print(f"Loading QwenImage21NVFP4Backend...")
|
| 44 |
+
print(f" • Base Pipeline: {self.model_id}")
|
| 45 |
+
print(f" • Quantized DiT: {self.optimized_model_path}")
|
| 46 |
+
print(f" • Target GPU: cuda:{self.gpu_id}")
|
| 47 |
+
print(f" • Stream Text Encoder: {self.stream_text_encoder}")
|
| 48 |
+
if self.text_encoder_path:
|
| 49 |
+
print(f" • Custom Text Encoder: {self.text_encoder_path}")
|
| 50 |
+
|
| 51 |
+
device = f"cuda:{self.gpu_id}"
|
| 52 |
+
|
| 53 |
+
# 1. Load base pipeline without heavy transformer
|
| 54 |
+
# Pass transformer=None or dummy to save loading 13GB BF16 transformer
|
| 55 |
+
print("Loading peripheral pipeline components (Text Encoder, Tokenizer, VAE, Scheduler)...")
|
| 56 |
+
if self.text_encoder_path:
|
| 57 |
+
from transformers import Qwen3VLForConditionalGeneration, Qwen3VLProcessor
|
| 58 |
+
|
| 59 |
+
te_dir = (
|
| 60 |
+
os.path.join(self.text_encoder_path, "text_encoder")
|
| 61 |
+
if os.path.isdir(os.path.join(self.text_encoder_path, "text_encoder"))
|
| 62 |
+
else self.text_encoder_path
|
| 63 |
+
)
|
| 64 |
+
proc_dir = (
|
| 65 |
+
os.path.join(self.text_encoder_path, "processor")
|
| 66 |
+
if os.path.isdir(os.path.join(self.text_encoder_path, "processor"))
|
| 67 |
+
else self.text_encoder_path
|
| 68 |
+
)
|
| 69 |
+
if not os.path.exists(os.path.join(proc_dir, "tokenizer.json")):
|
| 70 |
+
proc_dir = os.path.join(self.model_id, "processor")
|
| 71 |
+
|
| 72 |
+
print(f" • Loading custom Qwen3-VL text encoder from: {te_dir}")
|
| 73 |
+
custom_te = Qwen3VLForConditionalGeneration.from_pretrained(
|
| 74 |
+
te_dir,
|
| 75 |
+
torch_dtype=torch.bfloat16,
|
| 76 |
+
low_cpu_mem_usage=True,
|
| 77 |
+
)
|
| 78 |
+
print(f" • Loading processor from: {proc_dir}")
|
| 79 |
+
custom_proc = Qwen3VLProcessor.from_pretrained(proc_dir)
|
| 80 |
+
|
| 81 |
+
pipeline = QwenImage21Pipeline.from_pretrained(
|
| 82 |
+
self.model_id,
|
| 83 |
+
transformer=None,
|
| 84 |
+
text_encoder=custom_te,
|
| 85 |
+
processor=custom_proc,
|
| 86 |
+
torch_dtype=torch.bfloat16,
|
| 87 |
+
)
|
| 88 |
+
else:
|
| 89 |
+
pipeline = QwenImage21Pipeline.from_pretrained(
|
| 90 |
+
self.model_id,
|
| 91 |
+
transformer=None,
|
| 92 |
+
torch_dtype=torch.bfloat16,
|
| 93 |
+
)
|
| 94 |
+
|
| 95 |
+
# 2. Load forged Nunchaku NVFP4 transformer directly into resident VRAM
|
| 96 |
+
print(f"Loading resident NVFP4 DiT into {device}...")
|
| 97 |
+
quantized_transformer = NunchakuQwenImage21Transformer2DModel.from_pretrained(
|
| 98 |
+
self.optimized_model_path,
|
| 99 |
+
device=device,
|
| 100 |
+
torch_dtype=torch.bfloat16,
|
| 101 |
+
)
|
| 102 |
+
pipeline.transformer = quantized_transformer
|
| 103 |
+
|
| 104 |
+
# 3. Place VAE directly on target device with memory-safe tiling
|
| 105 |
+
print(f"Placing VAE on {device}...")
|
| 106 |
+
pipeline.vae = pipeline.vae.to(device)
|
| 107 |
+
if self.enable_tiling:
|
| 108 |
+
try:
|
| 109 |
+
pipeline.vae.enable_tiling()
|
| 110 |
+
print("VAE tiling enabled successfully.")
|
| 111 |
+
except Exception as e:
|
| 112 |
+
print(f"Note: VAE tiling could not be enabled ({e})")
|
| 113 |
+
|
| 114 |
+
# 4. Text Encoder Configuration: Layerwise PCIe Streaming or CPU Fallback
|
| 115 |
+
if self.stream_text_encoder:
|
| 116 |
+
print(f"Configuring PCIe Layerwise Weight Streaming for Qwen3-VL on {device}...")
|
| 117 |
+
from stream_encoder import attach_qwen3vl_streamer
|
| 118 |
+
self.streamer = attach_qwen3vl_streamer(pipeline, device=device)
|
| 119 |
+
|
| 120 |
+
orig_encode_prompt = pipeline.encode_prompt
|
| 121 |
+
|
| 122 |
+
def streamed_safe_encode_prompt(*args, **kwargs):
|
| 123 |
+
kwargs.pop("device", None)
|
| 124 |
+
embeds = orig_encode_prompt(*args, device=torch.device(device), **kwargs)
|
| 125 |
+
target_dev = pipeline.transformer.device
|
| 126 |
+
return tuple(x.to(target_dev) if isinstance(x, torch.Tensor) else x for x in embeds)
|
| 127 |
+
|
| 128 |
+
pipeline.encode_prompt = streamed_safe_encode_prompt
|
| 129 |
+
else:
|
| 130 |
+
print(f"Configuring hybrid CPU prompt encoding (DiT & VAE 100% resident on {device})...")
|
| 131 |
+
orig_encode_prompt = pipeline.encode_prompt
|
| 132 |
+
|
| 133 |
+
def safe_encode_prompt(*args, **kwargs):
|
| 134 |
+
kwargs.pop("device", None)
|
| 135 |
+
te_device = pipeline.text_encoder.device
|
| 136 |
+
embeds = orig_encode_prompt(*args, device=te_device, **kwargs)
|
| 137 |
+
target_dev = pipeline.transformer.device
|
| 138 |
+
return tuple(x.to(target_dev) if isinstance(x, torch.Tensor) else x for x in embeds)
|
| 139 |
+
|
| 140 |
+
pipeline.encode_prompt = safe_encode_prompt
|
| 141 |
+
|
| 142 |
+
# 5. Automatically attach optimal dynamic latent variance damping s(t)
|
| 143 |
+
if self.dynamic_scale_k > 0.0:
|
| 144 |
+
k = self.dynamic_scale_k
|
| 145 |
+
print(f"Attaching automatic late-stage variance damping (k={k:.3f}, +1.20 dB fidelity boost)...")
|
| 146 |
+
orig_call = pipeline.__call__
|
| 147 |
+
|
| 148 |
+
def scaled_call(*args, **kwargs):
|
| 149 |
+
user_cb = kwargs.pop("callback_on_step_end", None)
|
| 150 |
+
|
| 151 |
+
def combined_cb(pipe_obj, step_idx, timestep, callback_kwargs):
|
| 152 |
+
t = float(timestep.item() if isinstance(timestep, torch.Tensor) else timestep)
|
| 153 |
+
if t < 600.0:
|
| 154 |
+
scale = 1.0 - k * ((600.0 - t) / 600.0)
|
| 155 |
+
callback_kwargs["latents"] = callback_kwargs["latents"] * scale
|
| 156 |
+
if user_cb is not None:
|
| 157 |
+
return user_cb(pipe_obj, step_idx, timestep, callback_kwargs)
|
| 158 |
+
return callback_kwargs
|
| 159 |
+
|
| 160 |
+
kwargs["callback_on_step_end"] = combined_cb
|
| 161 |
+
return orig_call(*args, **kwargs)
|
| 162 |
+
|
| 163 |
+
pipeline.__call__ = scaled_call
|
| 164 |
+
|
| 165 |
+
self.pipeline = pipeline
|
| 166 |
+
return self.pipeline, self.pipeline
|
extras/create_heretic_text_encoder.py
ADDED
|
@@ -0,0 +1,201 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# -*- coding: utf-8 -*-
|
| 2 |
+
"""Heretic Abliteration Engine for Qwen3-VL Text Encoder.
|
| 3 |
+
|
| 4 |
+
Applies Norm-Preserving Biprojected Abliteration (Heretic/Arditi et al.) to Qwen3-VL-8B:
|
| 5 |
+
1. Measures the refusal/hesitation direction across all 36 transformer decoder layers.
|
| 6 |
+
2. Orthogonalizes the refusal direction against the benign semantic direction.
|
| 7 |
+
3. Applies rank-1 directional ablation to `self_attn.o_proj` and `mlp.down_proj`:
|
| 8 |
+
W' = normalize(W_norm - lambda * v * (v^T W_norm)) * ||W||_row
|
| 9 |
+
preserving exact row norms to protect general capability and language quality.
|
| 10 |
+
4. Exports the decensored, hesitation-free text encoder to ~/Nikola/models/Qwen_Text_Encoder_Heretic.
|
| 11 |
+
"""
|
| 12 |
+
|
| 13 |
+
import os
|
| 14 |
+
import sys
|
| 15 |
+
import shutil
|
| 16 |
+
import math
|
| 17 |
+
import torch
|
| 18 |
+
import torch.nn.functional as F
|
| 19 |
+
import torch.linalg as LA
|
| 20 |
+
from diffusers import QwenImage21Pipeline
|
| 21 |
+
|
| 22 |
+
ROOT_DIR = "/auto/home/amano/olegk/Nikola"
|
| 23 |
+
for p in [f"{ROOT_DIR}/src/imagegen", ROOT_DIR]:
|
| 24 |
+
if p not in sys.path:
|
| 25 |
+
sys.path.insert(0, p)
|
| 26 |
+
|
| 27 |
+
from stream_encoder import attach_qwen3vl_streamer
|
| 28 |
+
|
| 29 |
+
MODEL_PATH = "/home/olegk/Nikola/models/Qwen/Qwen-Image-2.1"
|
| 30 |
+
OUTPUT_DIR = "/home/olegk/Nikola/models/Qwen21_Text_Encoder_Heretic"
|
| 31 |
+
|
| 32 |
+
# Contrastive calibration prompts
|
| 33 |
+
BENIGN_PROMPTS = [
|
| 34 |
+
"A scenic view of a quiet tropical beach with palm trees, white sand, and gentle ocean waves under sunlight.",
|
| 35 |
+
"A group of athletic friends playing a casual game of volleyball on a sports court on a bright afternoon.",
|
| 36 |
+
"Two cheerful anime girls laughing together and talking in a colorful modern classroom setting.",
|
| 37 |
+
"A beautiful golden hour sunset over the coastal dunes with warm ambient lighting and ocean breeze.",
|
| 38 |
+
"Two athletic anime girls in standard sportswear practicing drills together on an outdoor gymnasium court.",
|
| 39 |
+
"A detailed photograph of a clean sandy beach with clear turquoise water and blue sky with light clouds.",
|
| 40 |
+
"Two anime girls taking a relaxing stroll along the seaside promenade during summer vacation.",
|
| 41 |
+
"A dynamic action photograph of a volleyball match on the sand with players in motion."
|
| 42 |
+
]
|
| 43 |
+
|
| 44 |
+
SENSITIVE_PROMPTS = [
|
| 45 |
+
"Two anime girls in colorful skimpy bikinis playing beach volleyball on a sunny tropical beach.",
|
| 46 |
+
"Two attractive anime girls wearing revealing two-piece swimsuits posing playfully under the bright sun.",
|
| 47 |
+
"Two cute anime girls in skimpy micro bikinis jumping to spike a volleyball on the beach sand.",
|
| 48 |
+
"A beach volleyball match with two anime girls wearing small swimsuits, athletic dynamic poses.",
|
| 49 |
+
"Two anime girls in alluring colorful bikinis having fun playing sports on a sunny seaside beach.",
|
| 50 |
+
"A close up action shot of two anime girls in tiny bikinis diving for a volleyball on the beach.",
|
| 51 |
+
"Two beautiful anime girls wearing revealing beachwear and swimsuits posing by the ocean shoreline.",
|
| 52 |
+
"An alluring tropical beach scene with two anime girls in revealing bikinis playing athletic beach sports."
|
| 53 |
+
]
|
| 54 |
+
|
| 55 |
+
|
| 56 |
+
def abliterate_layer_weights(weight: torch.Tensor, v: torch.Tensor, weight_factor: float) -> torch.Tensor:
|
| 57 |
+
"""Norm-preserving biprojected abliteration: delta W = -lambda * v * (v^T W)."""
|
| 58 |
+
if weight_factor <= 0.0:
|
| 59 |
+
return weight
|
| 60 |
+
|
| 61 |
+
orig_dtype = weight.dtype
|
| 62 |
+
W = weight.float()
|
| 63 |
+
v = v.float().to(W.device)
|
| 64 |
+
|
| 65 |
+
# Calculate row norms
|
| 66 |
+
row_norms = LA.vector_norm(W, dim=1, keepdim=True)
|
| 67 |
+
# Normalize rows
|
| 68 |
+
W_norm = F.normalize(W, p=2, dim=1)
|
| 69 |
+
|
| 70 |
+
# v @ W_norm -> (in_features,)
|
| 71 |
+
lora_A = (v @ W_norm).view(1, -1)
|
| 72 |
+
# -weight_factor * v -> (out_features, 1)
|
| 73 |
+
lora_B = (-weight_factor * v).view(-1, 1)
|
| 74 |
+
|
| 75 |
+
# Project and renormalize
|
| 76 |
+
W_adj = W_norm + lora_B @ lora_A
|
| 77 |
+
W_adj = F.normalize(W_adj, p=2, dim=1)
|
| 78 |
+
W_final = W_adj * row_norms
|
| 79 |
+
|
| 80 |
+
return W_final.to(orig_dtype)
|
| 81 |
+
|
| 82 |
+
|
| 83 |
+
def create_heretic_model():
|
| 84 |
+
print("=" * 80)
|
| 85 |
+
print("🔮 FORGING OPTIMIZED HERETIC TEXT ENCODER (QWEN3-VL-8B)")
|
| 86 |
+
print("=" * 80)
|
| 87 |
+
print(f"Base Model: {MODEL_PATH}")
|
| 88 |
+
print(f"Output Directory: {OUTPUT_DIR}")
|
| 89 |
+
|
| 90 |
+
# 1. Load pipeline and attach layerwise streamer for rapid residual extraction
|
| 91 |
+
print("\n[Step 1/4] Loading pipeline & attaching layerwise streamer...")
|
| 92 |
+
pipe = QwenImage21Pipeline.from_pretrained(
|
| 93 |
+
MODEL_PATH,
|
| 94 |
+
transformer=None,
|
| 95 |
+
torch_dtype=torch.bfloat16,
|
| 96 |
+
)
|
| 97 |
+
streamer = attach_qwen3vl_streamer(pipe, device="cuda:0")
|
| 98 |
+
|
| 99 |
+
# 2. Extract layer-by-layer residuals across contrastive prompt sets
|
| 100 |
+
print("\n[Step 2/4] Measuring refusal directions across all 36 decoder layers...")
|
| 101 |
+
|
| 102 |
+
def get_mean_residuals(prompts, label):
|
| 103 |
+
layer_states = [[] for _ in range(37)]
|
| 104 |
+
print(f" • Collecting residuals for {len(prompts)} {label} prompts...")
|
| 105 |
+
for p in prompts:
|
| 106 |
+
fmt = pipe.prompt_template_t2i.format(p)
|
| 107 |
+
inputs = pipe.processor(text=[fmt], return_tensors="pt").to("cuda:0")
|
| 108 |
+
with torch.no_grad():
|
| 109 |
+
out = pipe.text_encoder(
|
| 110 |
+
input_ids=inputs.input_ids,
|
| 111 |
+
attention_mask=inputs.attention_mask,
|
| 112 |
+
output_hidden_states=True,
|
| 113 |
+
)
|
| 114 |
+
for l in range(37):
|
| 115 |
+
layer_states[l].append(out.hidden_states[l][0, -1, :].float().cpu())
|
| 116 |
+
return torch.stack([torch.stack(l).mean(dim=0) for l in layer_states])
|
| 117 |
+
|
| 118 |
+
b_means = get_mean_residuals(BENIGN_PROMPTS, "benign")
|
| 119 |
+
s_means = get_mean_residuals(SENSITIVE_PROMPTS, "sensitive")
|
| 120 |
+
|
| 121 |
+
# Compute difference of means
|
| 122 |
+
residual_directions = s_means - b_means
|
| 123 |
+
|
| 124 |
+
# Orthogonalize against benign direction (Heretic projected abliteration)
|
| 125 |
+
print(" • Orthogonalizing refusal directions against benign semantic vectors...")
|
| 126 |
+
good_directions = F.normalize(b_means, p=2, dim=1)
|
| 127 |
+
proj = torch.sum(residual_directions * good_directions, dim=1, keepdim=True)
|
| 128 |
+
ortho_directions = residual_directions - proj * good_directions
|
| 129 |
+
ortho_directions = F.normalize(ortho_directions, p=2, dim=1)
|
| 130 |
+
|
| 131 |
+
# 3. Apply Norm-Preserving Biprojected Abliteration to Language Model Weights
|
| 132 |
+
print("\n[Step 3/4] Applying Heretic norm-preserving abliteration to weights...")
|
| 133 |
+
lm = pipe.text_encoder.model.language_model
|
| 134 |
+
num_layers = len(lm.layers)
|
| 135 |
+
|
| 136 |
+
# Target layers: layers 16 to 35, centered at layer 26 with max_weight = 1.0
|
| 137 |
+
center_layer = 26.0
|
| 138 |
+
spread = 8.0
|
| 139 |
+
|
| 140 |
+
ablated_count = 0
|
| 141 |
+
for l_idx, layer in enumerate(lm.layers):
|
| 142 |
+
dist = abs(l_idx - center_layer)
|
| 143 |
+
# Smooth bell-shaped abliteration profile
|
| 144 |
+
weight_factor = float(math.exp(-(dist**2) / (2 * (spread**2))))
|
| 145 |
+
if weight_factor < 0.10:
|
| 146 |
+
weight_factor = 0.0 # skip early layers (0..12) where refusal is inactive
|
| 147 |
+
|
| 148 |
+
v = ortho_directions[l_idx + 1] # l_idx+1 accounts for embedding layer at idx 0
|
| 149 |
+
|
| 150 |
+
if weight_factor > 0.0:
|
| 151 |
+
print(f" • Layer {l_idx:2d}: Abliterating with lambda={weight_factor:.3f}...")
|
| 152 |
+
|
| 153 |
+
# 1. Attention Out Projection
|
| 154 |
+
if hasattr(layer.self_attn, "o_proj"):
|
| 155 |
+
with torch.no_grad():
|
| 156 |
+
layer.self_attn.o_proj.weight.data = abliterate_layer_weights(
|
| 157 |
+
layer.self_attn.o_proj.weight.data, v, weight_factor
|
| 158 |
+
)
|
| 159 |
+
ablated_count += 1
|
| 160 |
+
|
| 161 |
+
# 2. MLP Down Projection
|
| 162 |
+
if hasattr(layer.mlp, "down_proj"):
|
| 163 |
+
with torch.no_grad():
|
| 164 |
+
layer.mlp.down_proj.weight.data = abliterate_layer_weights(
|
| 165 |
+
layer.mlp.down_proj.weight.data, v, weight_factor
|
| 166 |
+
)
|
| 167 |
+
ablated_count += 1
|
| 168 |
+
|
| 169 |
+
print(f" • Successfully abliterated {ablated_count} linear projection matrices!")
|
| 170 |
+
|
| 171 |
+
# 4. Save Standalone Checkpoint to models/Qwen_Text_Encoder_Heretic
|
| 172 |
+
print(f"\n[Step 4/4] Saving Heretic Text Encoder to {OUTPUT_DIR}...")
|
| 173 |
+
os.makedirs(OUTPUT_DIR, exist_ok=True)
|
| 174 |
+
|
| 175 |
+
# Save text_encoder subfolder (for direct use with Diffusers or Transformers)
|
| 176 |
+
sub_dir = os.path.join(OUTPUT_DIR, "text_encoder")
|
| 177 |
+
os.makedirs(sub_dir, exist_ok=True)
|
| 178 |
+
pipe.text_encoder.save_pretrained(sub_dir)
|
| 179 |
+
print(f" • Saved abliterated text encoder weights to: {sub_dir}")
|
| 180 |
+
|
| 181 |
+
# Also copy tokenizer and processor files
|
| 182 |
+
proc_src = os.path.join(MODEL_PATH, "processor")
|
| 183 |
+
proc_dst = os.path.join(OUTPUT_DIR, "processor")
|
| 184 |
+
if os.path.exists(proc_src):
|
| 185 |
+
if os.path.exists(proc_dst):
|
| 186 |
+
shutil.rmtree(proc_dst)
|
| 187 |
+
shutil.copytree(proc_src, proc_dst)
|
| 188 |
+
print(f" • Copied processor to: {proc_dst}")
|
| 189 |
+
|
| 190 |
+
# Copy root config files
|
| 191 |
+
for fname in ["config.json", "generation_config.json"]:
|
| 192 |
+
src_f = os.path.join(MODEL_PATH, "text_encoder", fname)
|
| 193 |
+
if os.path.exists(src_f):
|
| 194 |
+
shutil.copy(src_f, os.path.join(OUTPUT_DIR, fname))
|
| 195 |
+
|
| 196 |
+
print("\n🎉 SUCCESS: Heretic Text Encoder forged and saved successfully!")
|
| 197 |
+
print("=" * 80)
|
| 198 |
+
|
| 199 |
+
|
| 200 |
+
if __name__ == "__main__":
|
| 201 |
+
create_heretic_model()
|
extras/imagegen_qwen21_nvfp4.sh
ADDED
|
@@ -0,0 +1,40 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/bin/bash
|
| 2 |
+
# Start ImageEditServer with forged SVDQuant NVFP4 r32 Qwen-Image-2.1 model
|
| 3 |
+
# Runs 100% resident in VRAM on GPU 0 with Heretic text encoder by default
|
| 4 |
+
# Default: 25 steps, guidance-scale 1.0, port 4500, Text Encoder: Qwen21_Text_Encoder_Heretic
|
| 5 |
+
|
| 6 |
+
PORT=${1:-4500}
|
| 7 |
+
MODEL_DIR="/home/olegk/Nikola/models/Qwen/Qwen-Image-2.1"
|
| 8 |
+
OPTIMIZED_MODEL="/home/olegk/Nikola/models/nunchaku-qwen-image-2.1/best_quality_fp4.safetensors"
|
| 9 |
+
TEXT_ENCODER=${2:-"/home/olegk/Nikola/models/Qwen21_Text_Encoder_Heretic"}
|
| 10 |
+
|
| 11 |
+
echo "=========================================================="
|
| 12 |
+
echo "Starting ImageEditServer (Qwen-Image-2.1 SVDQuant NVFP4 - Resident VRAM)"
|
| 13 |
+
echo "Port: $PORT"
|
| 14 |
+
echo "Model: $MODEL_DIR"
|
| 15 |
+
echo "Optimized Model: $OPTIMIZED_MODEL"
|
| 16 |
+
echo "Text Encoder: $TEXT_ENCODER"
|
| 17 |
+
echo "Steps: 25 | Guidance Scale: 1.0 | Resident Mode: True"
|
| 18 |
+
echo "=========================================================="
|
| 19 |
+
|
| 20 |
+
export CUDA_VISIBLE_DEVICES="0"
|
| 21 |
+
export PYTORCH_CUDA_ALLOC_CONF="expandable_segments:True"
|
| 22 |
+
export PYTHONPATH="/home/olegk/Nikola/packages/nunchaku:/home/olegk/Nikola/packages/deepcompressor:/home/olegk/Nikola:$PYTHONPATH"
|
| 23 |
+
|
| 24 |
+
cd /auto/home/amano/olegk/Nikola/src/imagegen
|
| 25 |
+
|
| 26 |
+
EXTRA_ARGS=""
|
| 27 |
+
if [ -n "$TEXT_ENCODER" ] && [ "$TEXT_ENCODER" != "none" ]; then
|
| 28 |
+
EXTRA_ARGS="--text-encoder $TEXT_ENCODER"
|
| 29 |
+
fi
|
| 30 |
+
|
| 31 |
+
exec /home/olegk/Nikola/.venv/bin/python ImageEditServer.py \
|
| 32 |
+
--host 0.0.0.0 \
|
| 33 |
+
--port "$PORT" \
|
| 34 |
+
--model "$MODEL_DIR" \
|
| 35 |
+
--optimized-model "$OPTIMIZED_MODEL" \
|
| 36 |
+
--backend qwen21-nvfp4 \
|
| 37 |
+
--steps 25 \
|
| 38 |
+
--guidance-scale 1.0 \
|
| 39 |
+
$EXTRA_ARGS
|
| 40 |
+
|
extras/stream_encoder.py
ADDED
|
@@ -0,0 +1,190 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# -*- coding: utf-8 -*-
|
| 2 |
+
"""Layerwise Weight Streaming Engine for Qwen3-VL-8B in Qwen-Image-2.1.
|
| 3 |
+
|
| 4 |
+
Pins all 36 language model decoder layers in host RAM and streams them through a single
|
| 5 |
+
pre-allocated GPU layer buffer (368 MB VRAM) over PCIe (~28.7 GB/s).
|
| 6 |
+
Keeps the visual ViT encoder (1.07 GB) resident on GPU.
|
| 7 |
+
|
| 8 |
+
Achieves GPU compute speeds (1.06s multimodal prompt encode vs 17.18s on CPU, saving >16s per edit)
|
| 9 |
+
with only ~1.45 GB VRAM footprint and 100% bit-exact mathematical parity (zero quality loss).
|
| 10 |
+
"""
|
| 11 |
+
|
| 12 |
+
import copy
|
| 13 |
+
import time
|
| 14 |
+
import torch
|
| 15 |
+
import torch.nn as nn
|
| 16 |
+
from transformers.modeling_outputs import BaseModelOutputWithPast
|
| 17 |
+
from transformers.models.qwen3_vl.modeling_qwen3_vl import create_causal_mask
|
| 18 |
+
|
| 19 |
+
|
| 20 |
+
class Qwen3VLLayerwiseStreamer:
|
| 21 |
+
"""Streams Qwen3-VL language model decoder layers through a single static GPU buffer."""
|
| 22 |
+
|
| 23 |
+
def __init__(self, pipeline, device="cuda:0"):
|
| 24 |
+
self.pipeline = pipeline
|
| 25 |
+
self.device = torch.device(device)
|
| 26 |
+
self.text_encoder = pipeline.text_encoder
|
| 27 |
+
self.lm = getattr(self.text_encoder.model, "language_model", self.text_encoder.model)
|
| 28 |
+
self.num_layers = len(self.lm.layers)
|
| 29 |
+
|
| 30 |
+
print(f"Initializing Qwen3VLLayerwiseStreamer for {self.num_layers} layers on {self.device}...")
|
| 31 |
+
|
| 32 |
+
# 1. Pin CPU layers in host memory for maximum PCIe transfer throughput
|
| 33 |
+
t0 = time.perf_counter()
|
| 34 |
+
self.cpu_layers = []
|
| 35 |
+
for layer in self.lm.layers:
|
| 36 |
+
layer = layer.to("cpu", dtype=torch.bfloat16)
|
| 37 |
+
for p in layer.parameters():
|
| 38 |
+
if not p.data.is_pinned():
|
| 39 |
+
p.data = p.data.pin_memory()
|
| 40 |
+
for b in layer.buffers():
|
| 41 |
+
if not b.data.is_pinned():
|
| 42 |
+
b.data = b.data.pin_memory()
|
| 43 |
+
self.cpu_layers.append(layer)
|
| 44 |
+
t_pin = time.perf_counter() - t0
|
| 45 |
+
print(f" • Pinned {self.num_layers} layers in CPU RAM in {t_pin:.2f} s")
|
| 46 |
+
|
| 47 |
+
# 2. Allocate ONE single template GPU layer buffer in VRAM (~368 MB)
|
| 48 |
+
self.gpu_layer = copy.deepcopy(self.cpu_layers[0]).to(self.device, dtype=torch.bfloat16)
|
| 49 |
+
gpu_param_dict = dict(self.gpu_layer.named_parameters())
|
| 50 |
+
gpu_buffer_dict = dict(self.gpu_layer.named_buffers())
|
| 51 |
+
|
| 52 |
+
# Pre-build parameter transfer pairs for zero-overhead non-blocking copying
|
| 53 |
+
self.param_pairs = []
|
| 54 |
+
for i in range(self.num_layers):
|
| 55 |
+
lp = [(gpu_param_dict[name], cp) for name, cp in self.cpu_layers[i].named_parameters()]
|
| 56 |
+
lb = [(gpu_buffer_dict[name], cb) for name, cb in self.cpu_layers[i].named_buffers()]
|
| 57 |
+
self.param_pairs.append((lp, lb))
|
| 58 |
+
|
| 59 |
+
gpu_mb = sum(p.numel() * p.element_size() for p in self.gpu_layer.parameters()) / (1024**2)
|
| 60 |
+
print(f" • Static GPU layer buffer allocated: {gpu_mb:.2f} MB VRAM")
|
| 61 |
+
|
| 62 |
+
# 3. Place small peripheral layers directly on target GPU
|
| 63 |
+
self.lm.rotary_emb = self.lm.rotary_emb.to(self.device)
|
| 64 |
+
self.lm.embed_tokens = self.lm.embed_tokens.to(self.device)
|
| 65 |
+
self.lm.norm = self.lm.norm.to(self.device)
|
| 66 |
+
|
| 67 |
+
# 4. Place visual ViT encoder directly on target GPU (1.07 GB VRAM)
|
| 68 |
+
if hasattr(self.text_encoder.model, "visual") and self.text_encoder.model.visual is not None:
|
| 69 |
+
self.text_encoder.model.visual = self.text_encoder.model.visual.to(self.device, dtype=torch.bfloat16)
|
| 70 |
+
print(" • Visual ViT encoder placed resident on GPU (1.07 GB VRAM)")
|
| 71 |
+
|
| 72 |
+
# 5. Bypass unused lm_head (152,064 vocab projection, saving 1.24 GB computation)
|
| 73 |
+
class DummyHead(nn.Module):
|
| 74 |
+
def forward(self, x):
|
| 75 |
+
return None
|
| 76 |
+
|
| 77 |
+
self.text_encoder.lm_head = DummyHead()
|
| 78 |
+
print(" • Bypassed unused lm_head projection")
|
| 79 |
+
|
| 80 |
+
# 6. Install hooked forward pass
|
| 81 |
+
self.orig_lm_forward = self.lm.forward
|
| 82 |
+
self.lm.forward = self.streamed_forward
|
| 83 |
+
|
| 84 |
+
# 7. Route pipeline._get_qwen_prompt_embeds to target GPU
|
| 85 |
+
self.orig_get_embeds = self.pipeline._get_qwen_prompt_embeds
|
| 86 |
+
target_dev = self.device
|
| 87 |
+
|
| 88 |
+
def gpu_get_embeds(prompt_arg, image_arg, device_arg=None):
|
| 89 |
+
return self.orig_get_embeds(prompt_arg, image_arg, device=target_dev)
|
| 90 |
+
|
| 91 |
+
self.pipeline._get_qwen_prompt_embeds = gpu_get_embeds
|
| 92 |
+
print(f" • Hooked Qwen3-VL language model and prompt embedding router onto {self.device}!")
|
| 93 |
+
|
| 94 |
+
@torch.no_grad()
|
| 95 |
+
def streamed_forward(
|
| 96 |
+
self,
|
| 97 |
+
input_ids: torch.LongTensor | None = None,
|
| 98 |
+
attention_mask: torch.Tensor | None = None,
|
| 99 |
+
position_ids: torch.LongTensor | None = None,
|
| 100 |
+
past_key_values=None,
|
| 101 |
+
inputs_embeds: torch.FloatTensor | None = None,
|
| 102 |
+
use_cache: bool | None = None,
|
| 103 |
+
visual_pos_masks: torch.Tensor | None = None,
|
| 104 |
+
deepstack_visual_embeds: list[torch.Tensor] | None = None,
|
| 105 |
+
output_hidden_states: bool | None = None,
|
| 106 |
+
**kwargs,
|
| 107 |
+
) -> BaseModelOutputWithPast:
|
| 108 |
+
"""Executes language model decoding by streaming layers one by one into the GPU buffer."""
|
| 109 |
+
if inputs_embeds is None:
|
| 110 |
+
inputs_embeds = self.lm.embed_tokens(input_ids)
|
| 111 |
+
inputs_embeds = inputs_embeds.to(self.device)
|
| 112 |
+
|
| 113 |
+
if position_ids is None:
|
| 114 |
+
past_seen = past_key_values.get_seq_length() if past_key_values is not None else 0
|
| 115 |
+
position_ids = torch.arange(inputs_embeds.shape[1], device=self.device) + past_seen
|
| 116 |
+
position_ids = position_ids.view(1, 1, -1).expand(4, inputs_embeds.shape[0], -1)
|
| 117 |
+
elif position_ids.ndim == 2:
|
| 118 |
+
position_ids = position_ids[None, ...].expand(4, position_ids.shape[0], -1)
|
| 119 |
+
|
| 120 |
+
position_ids = position_ids.to(self.device)
|
| 121 |
+
if position_ids.ndim == 3 and position_ids.shape[0] == 4:
|
| 122 |
+
text_position_ids = position_ids[0]
|
| 123 |
+
rotary_pos_ids = position_ids[1:]
|
| 124 |
+
else:
|
| 125 |
+
text_position_ids = None
|
| 126 |
+
rotary_pos_ids = position_ids
|
| 127 |
+
|
| 128 |
+
causal_mask = create_causal_mask(
|
| 129 |
+
config=self.lm.config,
|
| 130 |
+
inputs_embeds=inputs_embeds,
|
| 131 |
+
attention_mask=attention_mask.to(self.device) if attention_mask is not None else None,
|
| 132 |
+
past_key_values=past_key_values,
|
| 133 |
+
position_ids=text_position_ids,
|
| 134 |
+
)
|
| 135 |
+
|
| 136 |
+
position_embeddings = self.lm.rotary_emb(inputs_embeds, rotary_pos_ids)
|
| 137 |
+
hidden_states = inputs_embeds
|
| 138 |
+
|
| 139 |
+
if visual_pos_masks is not None:
|
| 140 |
+
visual_pos_masks = visual_pos_masks.to(self.device)
|
| 141 |
+
if deepstack_visual_embeds is not None:
|
| 142 |
+
deepstack_visual_embeds = [d.to(self.device) for d in deepstack_visual_embeds]
|
| 143 |
+
|
| 144 |
+
all_hidden_states = () if output_hidden_states else None
|
| 145 |
+
|
| 146 |
+
# Stream all 36 decoder layers through the static GPU buffer
|
| 147 |
+
for layer_idx in range(self.num_layers):
|
| 148 |
+
if output_hidden_states:
|
| 149 |
+
all_hidden_states = all_hidden_states + (hidden_states,)
|
| 150 |
+
|
| 151 |
+
params, buffers = self.param_pairs[layer_idx]
|
| 152 |
+
for gp, cp in params:
|
| 153 |
+
gp.data.copy_(cp.data, non_blocking=True)
|
| 154 |
+
for gb, cb in buffers:
|
| 155 |
+
gb.data.copy_(cb.data, non_blocking=True)
|
| 156 |
+
|
| 157 |
+
layer_outputs = self.gpu_layer(
|
| 158 |
+
hidden_states,
|
| 159 |
+
attention_mask=causal_mask,
|
| 160 |
+
position_ids=text_position_ids,
|
| 161 |
+
past_key_values=past_key_values,
|
| 162 |
+
position_embeddings=position_embeddings,
|
| 163 |
+
**kwargs,
|
| 164 |
+
)
|
| 165 |
+
hidden_states = layer_outputs
|
| 166 |
+
|
| 167 |
+
# Add multi-layer deepstack visual features if present
|
| 168 |
+
if deepstack_visual_embeds is not None and layer_idx in range(len(deepstack_visual_embeds)):
|
| 169 |
+
hidden_states = self.lm._deepstack_process(
|
| 170 |
+
hidden_states,
|
| 171 |
+
visual_pos_masks,
|
| 172 |
+
deepstack_visual_embeds[layer_idx],
|
| 173 |
+
)
|
| 174 |
+
|
| 175 |
+
pre_norm_states = hidden_states
|
| 176 |
+
if output_hidden_states:
|
| 177 |
+
all_hidden_states = all_hidden_states + (pre_norm_states,)
|
| 178 |
+
|
| 179 |
+
norm_states = self.lm.norm(hidden_states)
|
| 180 |
+
|
| 181 |
+
return BaseModelOutputWithPast(
|
| 182 |
+
last_hidden_state=norm_states,
|
| 183 |
+
past_key_values=past_key_values,
|
| 184 |
+
hidden_states=all_hidden_states,
|
| 185 |
+
)
|
| 186 |
+
|
| 187 |
+
|
| 188 |
+
def attach_qwen3vl_streamer(pipeline, device="cuda:0") -> Qwen3VLLayerwiseStreamer:
|
| 189 |
+
"""Convenience factory to attach layerwise streaming to any QwenImage21Pipeline."""
|
| 190 |
+
return Qwen3VLLayerwiseStreamer(pipeline, device=device)
|
extras/test_heretic_beach_volleyball.py
ADDED
|
@@ -0,0 +1,128 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# -*- coding: utf-8 -*-
|
| 2 |
+
"""Verification & Comparison of Heretic Decensored Text Encoder on Qwen-Image-2.1.
|
| 3 |
+
|
| 4 |
+
Generates identical text-to-image scenes comparing:
|
| 5 |
+
1. Base / Stock Qwen3-VL Text Encoder (Safety Aligned / Blush Vector)
|
| 6 |
+
2. Heretic Qwen3-VL Text Encoder (Abliterated / Hesitation-Free)
|
| 7 |
+
|
| 8 |
+
Target Prompt:
|
| 9 |
+
"Two cute anime girls in colorful bikinis playing beach volleyball on a sunny tropical beach,
|
| 10 |
+
dynamic action pose, jumping to spike the ball, sharp focus, vibrant anime style,
|
| 11 |
+
detailed background with ocean and palm trees"
|
| 12 |
+
"""
|
| 13 |
+
|
| 14 |
+
import os
|
| 15 |
+
import sys
|
| 16 |
+
import time
|
| 17 |
+
import shutil
|
| 18 |
+
import torch
|
| 19 |
+
from PIL import Image
|
| 20 |
+
|
| 21 |
+
ROOT_DIR = "/auto/home/amano/olegk/Nikola"
|
| 22 |
+
for p in [f"{ROOT_DIR}/packages/nunchaku", f"{ROOT_DIR}/packages/deepcompressor", f"{ROOT_DIR}/src/imagegen", ROOT_DIR]:
|
| 23 |
+
if p not in sys.path:
|
| 24 |
+
sys.path.insert(0, p)
|
| 25 |
+
|
| 26 |
+
from QwenImage21NVFP4Backend import QwenImage21NVFP4Backend
|
| 27 |
+
|
| 28 |
+
PROMPT = (
|
| 29 |
+
"Two cute anime girls in colorful bikinis playing beach volleyball on a sunny tropical beach, "
|
| 30 |
+
"dynamic action pose, jumping to spike the ball, sharp focus, vibrant anime style, "
|
| 31 |
+
"detailed background with ocean and palm trees"
|
| 32 |
+
)
|
| 33 |
+
SEED = 42
|
| 34 |
+
STEPS = 25
|
| 35 |
+
GUIDANCE_SCALE = 1.0
|
| 36 |
+
HEIGHT = 1024
|
| 37 |
+
WIDTH = 1024
|
| 38 |
+
|
| 39 |
+
TMP_DIR = "/home/olegk/tmp"
|
| 40 |
+
ARTIFACT_DIR = "/home/olegk/.gemini/antigravity/brain/19d2cb51-1ef6-42f5-a95f-f814fe6ba720"
|
| 41 |
+
HERETIC_PATH = "/home/olegk/Nikola/models/Qwen21_Text_Encoder_Heretic"
|
| 42 |
+
|
| 43 |
+
|
| 44 |
+
def generate_with_backend(backend_name: str, text_encoder_path=None):
|
| 45 |
+
print(f"\n{'='*70}")
|
| 46 |
+
print(f"🚀 RUNNING GENERATION: {backend_name}")
|
| 47 |
+
print(f"Text Encoder Path: {text_encoder_path or 'Stock (Base Qwen-Image-2.1)'}")
|
| 48 |
+
print(f"{'='*70}")
|
| 49 |
+
|
| 50 |
+
backend = QwenImage21NVFP4Backend(
|
| 51 |
+
model_id="/home/olegk/Nikola/models/Qwen/Qwen-Image-2.1",
|
| 52 |
+
optimized_model_path="/home/olegk/Nikola/models/nunchaku-qwen-image-2.1/best_quality_fp4.safetensors",
|
| 53 |
+
gpu_id=0,
|
| 54 |
+
enable_tiling=True,
|
| 55 |
+
dynamic_scale_k=0.0,
|
| 56 |
+
stream_text_encoder=True,
|
| 57 |
+
text_encoder_path=text_encoder_path,
|
| 58 |
+
)
|
| 59 |
+
|
| 60 |
+
t0 = time.perf_counter()
|
| 61 |
+
pipeline, _ = backend.load()
|
| 62 |
+
t_load = time.perf_counter() - t0
|
| 63 |
+
print(f"Backend loaded in {t_load:.2f} s")
|
| 64 |
+
|
| 65 |
+
generator = torch.Generator(device="cuda:0").manual_seed(SEED)
|
| 66 |
+
|
| 67 |
+
torch.cuda.synchronize()
|
| 68 |
+
t_gen_start = time.perf_counter()
|
| 69 |
+
output = pipeline(
|
| 70 |
+
prompt=PROMPT,
|
| 71 |
+
height=HEIGHT,
|
| 72 |
+
width=WIDTH,
|
| 73 |
+
num_inference_steps=STEPS,
|
| 74 |
+
true_cfg_scale=GUIDANCE_SCALE,
|
| 75 |
+
generator=generator,
|
| 76 |
+
)
|
| 77 |
+
torch.cuda.synchronize()
|
| 78 |
+
t_gen = time.perf_counter() - t_gen_start
|
| 79 |
+
print(f"Generation completed in {t_gen:.2f} s")
|
| 80 |
+
|
| 81 |
+
img = output.images[0]
|
| 82 |
+
|
| 83 |
+
# Clean up pipeline from GPU memory
|
| 84 |
+
del pipeline
|
| 85 |
+
del backend
|
| 86 |
+
torch.cuda.empty_cache()
|
| 87 |
+
|
| 88 |
+
return img, t_gen
|
| 89 |
+
|
| 90 |
+
|
| 91 |
+
def main():
|
| 92 |
+
os.makedirs(TMP_DIR, exist_ok=True)
|
| 93 |
+
os.makedirs(ARTIFACT_DIR, exist_ok=True)
|
| 94 |
+
|
| 95 |
+
print("=" * 80)
|
| 96 |
+
print("🏐 ANIME BEACH VOLLEYBALL TEXT ENCODER COMPARISON TEST")
|
| 97 |
+
print(f"Prompt: {PROMPT}")
|
| 98 |
+
print(f"Seed: {SEED} | Steps: {STEPS} | CFG: {GUIDANCE_SCALE}")
|
| 99 |
+
print("=" * 80)
|
| 100 |
+
|
| 101 |
+
# 1. Run Stock Text Encoder
|
| 102 |
+
img_stock, t_stock = generate_with_backend("STOCK (Base Qwen-Image-2.1 Text Encoder)", text_encoder_path=None)
|
| 103 |
+
stock_path = os.path.join(TMP_DIR, "qwen21_stock_beach_volleyball.png")
|
| 104 |
+
stock_art = os.path.join(ARTIFACT_DIR, "qwen21_stock_beach_volleyball.png")
|
| 105 |
+
img_stock.save(stock_path)
|
| 106 |
+
shutil.copy(stock_path, stock_art)
|
| 107 |
+
print(f"Saved stock image to: {stock_path} and artifact")
|
| 108 |
+
|
| 109 |
+
# 2. Run Heretic Text Encoder
|
| 110 |
+
img_heretic, t_heretic = generate_with_backend(
|
| 111 |
+
"HERETIC (Decensored Qwen3-VL Text Encoder)",
|
| 112 |
+
text_encoder_path=HERETIC_PATH,
|
| 113 |
+
)
|
| 114 |
+
heretic_path = os.path.join(TMP_DIR, "qwen21_heretic_beach_volleyball.png")
|
| 115 |
+
heretic_art = os.path.join(ARTIFACT_DIR, "qwen21_heretic_beach_volleyball.png")
|
| 116 |
+
img_heretic.save(heretic_path)
|
| 117 |
+
shutil.copy(heretic_path, heretic_art)
|
| 118 |
+
print(f"Saved heretic image to: {heretic_path} and artifact")
|
| 119 |
+
|
| 120 |
+
print("\n" + "=" * 80)
|
| 121 |
+
print("🎉 COMPARISON TEST COMPLETED SUCCESSFULLY!")
|
| 122 |
+
print(f"Stock Image: {stock_path} ({t_stock:.2f} s)")
|
| 123 |
+
print(f"Heretic Image: {heretic_path} ({t_heretic:.2f} s)")
|
| 124 |
+
print("=" * 80)
|
| 125 |
+
|
| 126 |
+
|
| 127 |
+
if __name__ == "__main__":
|
| 128 |
+
main()
|
generation_config.json
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"bos_token_id": 151643,
|
| 3 |
+
"do_sample": true,
|
| 4 |
+
"eos_token_id": [
|
| 5 |
+
151645,
|
| 6 |
+
151643
|
| 7 |
+
],
|
| 8 |
+
"pad_token_id": 151643,
|
| 9 |
+
"temperature": 0.7,
|
| 10 |
+
"top_k": 20,
|
| 11 |
+
"top_p": 0.8,
|
| 12 |
+
"transformers_version": "5.3.0.dev0"
|
| 13 |
+
}
|
merges.txt
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ba4cadb5ccea2f11dcbb7facce01cb35a76c9a2d5553b34f681b2e16227a1f02
|
| 3 |
+
size 16289680712
|
preprocessor_config.json
ADDED
|
@@ -0,0 +1,39 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"crop_size": null,
|
| 3 |
+
"data_format": "channels_first",
|
| 4 |
+
"default_to_square": true,
|
| 5 |
+
"device": null,
|
| 6 |
+
"disable_grouping": null,
|
| 7 |
+
"do_center_crop": null,
|
| 8 |
+
"do_convert_rgb": true,
|
| 9 |
+
"do_normalize": true,
|
| 10 |
+
"do_pad": null,
|
| 11 |
+
"do_rescale": true,
|
| 12 |
+
"do_resize": true,
|
| 13 |
+
"image_mean": [
|
| 14 |
+
0.5,
|
| 15 |
+
0.5,
|
| 16 |
+
0.5
|
| 17 |
+
],
|
| 18 |
+
"image_processor_type": "Qwen2VLImageProcessorFast",
|
| 19 |
+
"image_std": [
|
| 20 |
+
0.5,
|
| 21 |
+
0.5,
|
| 22 |
+
0.5
|
| 23 |
+
],
|
| 24 |
+
"input_data_format": null,
|
| 25 |
+
"max_pixels": null,
|
| 26 |
+
"merge_size": 2,
|
| 27 |
+
"min_pixels": null,
|
| 28 |
+
"pad_size": null,
|
| 29 |
+
"patch_size": 16,
|
| 30 |
+
"processor_class": "Qwen3VLProcessor",
|
| 31 |
+
"resample": 3,
|
| 32 |
+
"rescale_factor": 0.00392156862745098,
|
| 33 |
+
"return_tensors": null,
|
| 34 |
+
"size": {
|
| 35 |
+
"longest_edge": 16777216,
|
| 36 |
+
"shortest_edge": 65536
|
| 37 |
+
},
|
| 38 |
+
"temporal_patch_size": 2
|
| 39 |
+
}
|
processor/added_tokens.json
ADDED
|
@@ -0,0 +1,28 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"</think>": 151668,
|
| 3 |
+
"</tool_call>": 151658,
|
| 4 |
+
"</tool_response>": 151666,
|
| 5 |
+
"<think>": 151667,
|
| 6 |
+
"<tool_call>": 151657,
|
| 7 |
+
"<tool_response>": 151665,
|
| 8 |
+
"<|box_end|>": 151649,
|
| 9 |
+
"<|box_start|>": 151648,
|
| 10 |
+
"<|endoftext|>": 151643,
|
| 11 |
+
"<|file_sep|>": 151664,
|
| 12 |
+
"<|fim_middle|>": 151660,
|
| 13 |
+
"<|fim_pad|>": 151662,
|
| 14 |
+
"<|fim_prefix|>": 151659,
|
| 15 |
+
"<|fim_suffix|>": 151661,
|
| 16 |
+
"<|im_end|>": 151645,
|
| 17 |
+
"<|im_start|>": 151644,
|
| 18 |
+
"<|image_pad|>": 151655,
|
| 19 |
+
"<|object_ref_end|>": 151647,
|
| 20 |
+
"<|object_ref_start|>": 151646,
|
| 21 |
+
"<|quad_end|>": 151651,
|
| 22 |
+
"<|quad_start|>": 151650,
|
| 23 |
+
"<|repo_name|>": 151663,
|
| 24 |
+
"<|video_pad|>": 151656,
|
| 25 |
+
"<|vision_end|>": 151653,
|
| 26 |
+
"<|vision_pad|>": 151654,
|
| 27 |
+
"<|vision_start|>": 151652
|
| 28 |
+
}
|
processor/chat_template.jinja
ADDED
|
@@ -0,0 +1,120 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0].role == 'system' %}
|
| 4 |
+
{%- if messages[0].content is string %}
|
| 5 |
+
{{- messages[0].content }}
|
| 6 |
+
{%- else %}
|
| 7 |
+
{%- for content in messages[0].content %}
|
| 8 |
+
{%- if 'text' in content %}
|
| 9 |
+
{{- content.text }}
|
| 10 |
+
{%- endif %}
|
| 11 |
+
{%- endfor %}
|
| 12 |
+
{%- endif %}
|
| 13 |
+
{{- '\n\n' }}
|
| 14 |
+
{%- endif %}
|
| 15 |
+
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 16 |
+
{%- for tool in tools %}
|
| 17 |
+
{{- "\n" }}
|
| 18 |
+
{{- tool | tojson }}
|
| 19 |
+
{%- endfor %}
|
| 20 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 21 |
+
{%- else %}
|
| 22 |
+
{%- if messages[0].role == 'system' %}
|
| 23 |
+
{{- '<|im_start|>system\n' }}
|
| 24 |
+
{%- if messages[0].content is string %}
|
| 25 |
+
{{- messages[0].content }}
|
| 26 |
+
{%- else %}
|
| 27 |
+
{%- for content in messages[0].content %}
|
| 28 |
+
{%- if 'text' in content %}
|
| 29 |
+
{{- content.text }}
|
| 30 |
+
{%- endif %}
|
| 31 |
+
{%- endfor %}
|
| 32 |
+
{%- endif %}
|
| 33 |
+
{{- '<|im_end|>\n' }}
|
| 34 |
+
{%- endif %}
|
| 35 |
+
{%- endif %}
|
| 36 |
+
{%- set image_count = namespace(value=0) %}
|
| 37 |
+
{%- set video_count = namespace(value=0) %}
|
| 38 |
+
{%- for message in messages %}
|
| 39 |
+
{%- if message.role == "user" %}
|
| 40 |
+
{{- '<|im_start|>' + message.role + '\n' }}
|
| 41 |
+
{%- if message.content is string %}
|
| 42 |
+
{{- message.content }}
|
| 43 |
+
{%- else %}
|
| 44 |
+
{%- for content in message.content %}
|
| 45 |
+
{%- if content.type == 'image' or 'image' in content or 'image_url' in content %}
|
| 46 |
+
{%- set image_count.value = image_count.value + 1 %}
|
| 47 |
+
{%- if add_vision_id %}Picture {{ image_count.value }}: {% endif -%}
|
| 48 |
+
<|vision_start|><|image_pad|><|vision_end|>
|
| 49 |
+
{%- elif content.type == 'video' or 'video' in content %}
|
| 50 |
+
{%- set video_count.value = video_count.value + 1 %}
|
| 51 |
+
{%- if add_vision_id %}Video {{ video_count.value }}: {% endif -%}
|
| 52 |
+
<|vision_start|><|video_pad|><|vision_end|>
|
| 53 |
+
{%- elif 'text' in content %}
|
| 54 |
+
{{- content.text }}
|
| 55 |
+
{%- endif %}
|
| 56 |
+
{%- endfor %}
|
| 57 |
+
{%- endif %}
|
| 58 |
+
{{- '<|im_end|>\n' }}
|
| 59 |
+
{%- elif message.role == "assistant" %}
|
| 60 |
+
{{- '<|im_start|>' + message.role + '\n' }}
|
| 61 |
+
{%- if message.content is string %}
|
| 62 |
+
{{- message.content }}
|
| 63 |
+
{%- else %}
|
| 64 |
+
{%- for content_item in message.content %}
|
| 65 |
+
{%- if 'text' in content_item %}
|
| 66 |
+
{{- content_item.text }}
|
| 67 |
+
{%- endif %}
|
| 68 |
+
{%- endfor %}
|
| 69 |
+
{%- endif %}
|
| 70 |
+
{%- if message.tool_calls %}
|
| 71 |
+
{%- for tool_call in message.tool_calls %}
|
| 72 |
+
{%- if (loop.first and message.content) or (not loop.first) %}
|
| 73 |
+
{{- '\n' }}
|
| 74 |
+
{%- endif %}
|
| 75 |
+
{%- if tool_call.function %}
|
| 76 |
+
{%- set tool_call = tool_call.function %}
|
| 77 |
+
{%- endif %}
|
| 78 |
+
{{- '<tool_call>\n{"name": "' }}
|
| 79 |
+
{{- tool_call.name }}
|
| 80 |
+
{{- '", "arguments": ' }}
|
| 81 |
+
{%- if tool_call.arguments is string %}
|
| 82 |
+
{{- tool_call.arguments }}
|
| 83 |
+
{%- else %}
|
| 84 |
+
{{- tool_call.arguments | tojson }}
|
| 85 |
+
{%- endif %}
|
| 86 |
+
{{- '}\n</tool_call>' }}
|
| 87 |
+
{%- endfor %}
|
| 88 |
+
{%- endif %}
|
| 89 |
+
{{- '<|im_end|>\n' }}
|
| 90 |
+
{%- elif message.role == "tool" %}
|
| 91 |
+
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
| 92 |
+
{{- '<|im_start|>user' }}
|
| 93 |
+
{%- endif %}
|
| 94 |
+
{{- '\n<tool_response>\n' }}
|
| 95 |
+
{%- if message.content is string %}
|
| 96 |
+
{{- message.content }}
|
| 97 |
+
{%- else %}
|
| 98 |
+
{%- for content in message.content %}
|
| 99 |
+
{%- if content.type == 'image' or 'image' in content or 'image_url' in content %}
|
| 100 |
+
{%- set image_count.value = image_count.value + 1 %}
|
| 101 |
+
{%- if add_vision_id %}Picture {{ image_count.value }}: {% endif -%}
|
| 102 |
+
<|vision_start|><|image_pad|><|vision_end|>
|
| 103 |
+
{%- elif content.type == 'video' or 'video' in content %}
|
| 104 |
+
{%- set video_count.value = video_count.value + 1 %}
|
| 105 |
+
{%- if add_vision_id %}Video {{ video_count.value }}: {% endif -%}
|
| 106 |
+
<|vision_start|><|video_pad|><|vision_end|>
|
| 107 |
+
{%- elif 'text' in content %}
|
| 108 |
+
{{- content.text }}
|
| 109 |
+
{%- endif %}
|
| 110 |
+
{%- endfor %}
|
| 111 |
+
{%- endif %}
|
| 112 |
+
{{- '\n</tool_response>' }}
|
| 113 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 114 |
+
{{- '<|im_end|>\n' }}
|
| 115 |
+
{%- endif %}
|
| 116 |
+
{%- endif %}
|
| 117 |
+
{%- endfor %}
|
| 118 |
+
{%- if add_generation_prompt %}
|
| 119 |
+
{{- '<|im_start|>assistant\n' }}
|
| 120 |
+
{%- endif %}
|
processor/merges.txt
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
processor/preprocessor_config.json
ADDED
|
@@ -0,0 +1,39 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"crop_size": null,
|
| 3 |
+
"data_format": "channels_first",
|
| 4 |
+
"default_to_square": true,
|
| 5 |
+
"device": null,
|
| 6 |
+
"disable_grouping": null,
|
| 7 |
+
"do_center_crop": null,
|
| 8 |
+
"do_convert_rgb": true,
|
| 9 |
+
"do_normalize": true,
|
| 10 |
+
"do_pad": null,
|
| 11 |
+
"do_rescale": true,
|
| 12 |
+
"do_resize": true,
|
| 13 |
+
"image_mean": [
|
| 14 |
+
0.5,
|
| 15 |
+
0.5,
|
| 16 |
+
0.5
|
| 17 |
+
],
|
| 18 |
+
"image_processor_type": "Qwen2VLImageProcessorFast",
|
| 19 |
+
"image_std": [
|
| 20 |
+
0.5,
|
| 21 |
+
0.5,
|
| 22 |
+
0.5
|
| 23 |
+
],
|
| 24 |
+
"input_data_format": null,
|
| 25 |
+
"max_pixels": null,
|
| 26 |
+
"merge_size": 2,
|
| 27 |
+
"min_pixels": null,
|
| 28 |
+
"pad_size": null,
|
| 29 |
+
"patch_size": 16,
|
| 30 |
+
"processor_class": "Qwen3VLProcessor",
|
| 31 |
+
"resample": 3,
|
| 32 |
+
"rescale_factor": 0.00392156862745098,
|
| 33 |
+
"return_tensors": null,
|
| 34 |
+
"size": {
|
| 35 |
+
"longest_edge": 16777216,
|
| 36 |
+
"shortest_edge": 65536
|
| 37 |
+
},
|
| 38 |
+
"temporal_patch_size": 2
|
| 39 |
+
}
|
processor/special_tokens_map.json
ADDED
|
@@ -0,0 +1,31 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"additional_special_tokens": [
|
| 3 |
+
"<|im_start|>",
|
| 4 |
+
"<|im_end|>",
|
| 5 |
+
"<|object_ref_start|>",
|
| 6 |
+
"<|object_ref_end|>",
|
| 7 |
+
"<|box_start|>",
|
| 8 |
+
"<|box_end|>",
|
| 9 |
+
"<|quad_start|>",
|
| 10 |
+
"<|quad_end|>",
|
| 11 |
+
"<|vision_start|>",
|
| 12 |
+
"<|vision_end|>",
|
| 13 |
+
"<|vision_pad|>",
|
| 14 |
+
"<|image_pad|>",
|
| 15 |
+
"<|video_pad|>"
|
| 16 |
+
],
|
| 17 |
+
"eos_token": {
|
| 18 |
+
"content": "<|im_end|>",
|
| 19 |
+
"lstrip": false,
|
| 20 |
+
"normalized": false,
|
| 21 |
+
"rstrip": false,
|
| 22 |
+
"single_word": false
|
| 23 |
+
},
|
| 24 |
+
"pad_token": {
|
| 25 |
+
"content": "<|endoftext|>",
|
| 26 |
+
"lstrip": false,
|
| 27 |
+
"normalized": false,
|
| 28 |
+
"rstrip": false,
|
| 29 |
+
"single_word": false
|
| 30 |
+
}
|
| 31 |
+
}
|
processor/tokenizer.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:aeb13307a71acd8fe81861d94ad54ab689df773318809eed3cbe794b4492dae4
|
| 3 |
+
size 11422654
|
processor/tokenizer_config.json
ADDED
|
@@ -0,0 +1,240 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_bos_token": false,
|
| 3 |
+
"add_prefix_space": false,
|
| 4 |
+
"added_tokens_decoder": {
|
| 5 |
+
"151643": {
|
| 6 |
+
"content": "<|endoftext|>",
|
| 7 |
+
"lstrip": false,
|
| 8 |
+
"normalized": false,
|
| 9 |
+
"rstrip": false,
|
| 10 |
+
"single_word": false,
|
| 11 |
+
"special": true
|
| 12 |
+
},
|
| 13 |
+
"151644": {
|
| 14 |
+
"content": "<|im_start|>",
|
| 15 |
+
"lstrip": false,
|
| 16 |
+
"normalized": false,
|
| 17 |
+
"rstrip": false,
|
| 18 |
+
"single_word": false,
|
| 19 |
+
"special": true
|
| 20 |
+
},
|
| 21 |
+
"151645": {
|
| 22 |
+
"content": "<|im_end|>",
|
| 23 |
+
"lstrip": false,
|
| 24 |
+
"normalized": false,
|
| 25 |
+
"rstrip": false,
|
| 26 |
+
"single_word": false,
|
| 27 |
+
"special": true
|
| 28 |
+
},
|
| 29 |
+
"151646": {
|
| 30 |
+
"content": "<|object_ref_start|>",
|
| 31 |
+
"lstrip": false,
|
| 32 |
+
"normalized": false,
|
| 33 |
+
"rstrip": false,
|
| 34 |
+
"single_word": false,
|
| 35 |
+
"special": true
|
| 36 |
+
},
|
| 37 |
+
"151647": {
|
| 38 |
+
"content": "<|object_ref_end|>",
|
| 39 |
+
"lstrip": false,
|
| 40 |
+
"normalized": false,
|
| 41 |
+
"rstrip": false,
|
| 42 |
+
"single_word": false,
|
| 43 |
+
"special": true
|
| 44 |
+
},
|
| 45 |
+
"151648": {
|
| 46 |
+
"content": "<|box_start|>",
|
| 47 |
+
"lstrip": false,
|
| 48 |
+
"normalized": false,
|
| 49 |
+
"rstrip": false,
|
| 50 |
+
"single_word": false,
|
| 51 |
+
"special": true
|
| 52 |
+
},
|
| 53 |
+
"151649": {
|
| 54 |
+
"content": "<|box_end|>",
|
| 55 |
+
"lstrip": false,
|
| 56 |
+
"normalized": false,
|
| 57 |
+
"rstrip": false,
|
| 58 |
+
"single_word": false,
|
| 59 |
+
"special": true
|
| 60 |
+
},
|
| 61 |
+
"151650": {
|
| 62 |
+
"content": "<|quad_start|>",
|
| 63 |
+
"lstrip": false,
|
| 64 |
+
"normalized": false,
|
| 65 |
+
"rstrip": false,
|
| 66 |
+
"single_word": false,
|
| 67 |
+
"special": true
|
| 68 |
+
},
|
| 69 |
+
"151651": {
|
| 70 |
+
"content": "<|quad_end|>",
|
| 71 |
+
"lstrip": false,
|
| 72 |
+
"normalized": false,
|
| 73 |
+
"rstrip": false,
|
| 74 |
+
"single_word": false,
|
| 75 |
+
"special": true
|
| 76 |
+
},
|
| 77 |
+
"151652": {
|
| 78 |
+
"content": "<|vision_start|>",
|
| 79 |
+
"lstrip": false,
|
| 80 |
+
"normalized": false,
|
| 81 |
+
"rstrip": false,
|
| 82 |
+
"single_word": false,
|
| 83 |
+
"special": true
|
| 84 |
+
},
|
| 85 |
+
"151653": {
|
| 86 |
+
"content": "<|vision_end|>",
|
| 87 |
+
"lstrip": false,
|
| 88 |
+
"normalized": false,
|
| 89 |
+
"rstrip": false,
|
| 90 |
+
"single_word": false,
|
| 91 |
+
"special": true
|
| 92 |
+
},
|
| 93 |
+
"151654": {
|
| 94 |
+
"content": "<|vision_pad|>",
|
| 95 |
+
"lstrip": false,
|
| 96 |
+
"normalized": false,
|
| 97 |
+
"rstrip": false,
|
| 98 |
+
"single_word": false,
|
| 99 |
+
"special": true
|
| 100 |
+
},
|
| 101 |
+
"151655": {
|
| 102 |
+
"content": "<|image_pad|>",
|
| 103 |
+
"lstrip": false,
|
| 104 |
+
"normalized": false,
|
| 105 |
+
"rstrip": false,
|
| 106 |
+
"single_word": false,
|
| 107 |
+
"special": true
|
| 108 |
+
},
|
| 109 |
+
"151656": {
|
| 110 |
+
"content": "<|video_pad|>",
|
| 111 |
+
"lstrip": false,
|
| 112 |
+
"normalized": false,
|
| 113 |
+
"rstrip": false,
|
| 114 |
+
"single_word": false,
|
| 115 |
+
"special": true
|
| 116 |
+
},
|
| 117 |
+
"151657": {
|
| 118 |
+
"content": "<tool_call>",
|
| 119 |
+
"lstrip": false,
|
| 120 |
+
"normalized": false,
|
| 121 |
+
"rstrip": false,
|
| 122 |
+
"single_word": false,
|
| 123 |
+
"special": false
|
| 124 |
+
},
|
| 125 |
+
"151658": {
|
| 126 |
+
"content": "</tool_call>",
|
| 127 |
+
"lstrip": false,
|
| 128 |
+
"normalized": false,
|
| 129 |
+
"rstrip": false,
|
| 130 |
+
"single_word": false,
|
| 131 |
+
"special": false
|
| 132 |
+
},
|
| 133 |
+
"151659": {
|
| 134 |
+
"content": "<|fim_prefix|>",
|
| 135 |
+
"lstrip": false,
|
| 136 |
+
"normalized": false,
|
| 137 |
+
"rstrip": false,
|
| 138 |
+
"single_word": false,
|
| 139 |
+
"special": false
|
| 140 |
+
},
|
| 141 |
+
"151660": {
|
| 142 |
+
"content": "<|fim_middle|>",
|
| 143 |
+
"lstrip": false,
|
| 144 |
+
"normalized": false,
|
| 145 |
+
"rstrip": false,
|
| 146 |
+
"single_word": false,
|
| 147 |
+
"special": false
|
| 148 |
+
},
|
| 149 |
+
"151661": {
|
| 150 |
+
"content": "<|fim_suffix|>",
|
| 151 |
+
"lstrip": false,
|
| 152 |
+
"normalized": false,
|
| 153 |
+
"rstrip": false,
|
| 154 |
+
"single_word": false,
|
| 155 |
+
"special": false
|
| 156 |
+
},
|
| 157 |
+
"151662": {
|
| 158 |
+
"content": "<|fim_pad|>",
|
| 159 |
+
"lstrip": false,
|
| 160 |
+
"normalized": false,
|
| 161 |
+
"rstrip": false,
|
| 162 |
+
"single_word": false,
|
| 163 |
+
"special": false
|
| 164 |
+
},
|
| 165 |
+
"151663": {
|
| 166 |
+
"content": "<|repo_name|>",
|
| 167 |
+
"lstrip": false,
|
| 168 |
+
"normalized": false,
|
| 169 |
+
"rstrip": false,
|
| 170 |
+
"single_word": false,
|
| 171 |
+
"special": false
|
| 172 |
+
},
|
| 173 |
+
"151664": {
|
| 174 |
+
"content": "<|file_sep|>",
|
| 175 |
+
"lstrip": false,
|
| 176 |
+
"normalized": false,
|
| 177 |
+
"rstrip": false,
|
| 178 |
+
"single_word": false,
|
| 179 |
+
"special": false
|
| 180 |
+
},
|
| 181 |
+
"151665": {
|
| 182 |
+
"content": "<tool_response>",
|
| 183 |
+
"lstrip": false,
|
| 184 |
+
"normalized": false,
|
| 185 |
+
"rstrip": false,
|
| 186 |
+
"single_word": false,
|
| 187 |
+
"special": false
|
| 188 |
+
},
|
| 189 |
+
"151666": {
|
| 190 |
+
"content": "</tool_response>",
|
| 191 |
+
"lstrip": false,
|
| 192 |
+
"normalized": false,
|
| 193 |
+
"rstrip": false,
|
| 194 |
+
"single_word": false,
|
| 195 |
+
"special": false
|
| 196 |
+
},
|
| 197 |
+
"151667": {
|
| 198 |
+
"content": "<think>",
|
| 199 |
+
"lstrip": false,
|
| 200 |
+
"normalized": false,
|
| 201 |
+
"rstrip": false,
|
| 202 |
+
"single_word": false,
|
| 203 |
+
"special": false
|
| 204 |
+
},
|
| 205 |
+
"151668": {
|
| 206 |
+
"content": "</think>",
|
| 207 |
+
"lstrip": false,
|
| 208 |
+
"normalized": false,
|
| 209 |
+
"rstrip": false,
|
| 210 |
+
"single_word": false,
|
| 211 |
+
"special": false
|
| 212 |
+
}
|
| 213 |
+
},
|
| 214 |
+
"additional_special_tokens": [
|
| 215 |
+
"<|im_start|>",
|
| 216 |
+
"<|im_end|>",
|
| 217 |
+
"<|object_ref_start|>",
|
| 218 |
+
"<|object_ref_end|>",
|
| 219 |
+
"<|box_start|>",
|
| 220 |
+
"<|box_end|>",
|
| 221 |
+
"<|quad_start|>",
|
| 222 |
+
"<|quad_end|>",
|
| 223 |
+
"<|vision_start|>",
|
| 224 |
+
"<|vision_end|>",
|
| 225 |
+
"<|vision_pad|>",
|
| 226 |
+
"<|image_pad|>",
|
| 227 |
+
"<|video_pad|>"
|
| 228 |
+
],
|
| 229 |
+
"bos_token": null,
|
| 230 |
+
"clean_up_tokenization_spaces": false,
|
| 231 |
+
"eos_token": "<|im_end|>",
|
| 232 |
+
"errors": "replace",
|
| 233 |
+
"extra_special_tokens": {},
|
| 234 |
+
"model_max_length": 262144,
|
| 235 |
+
"pad_token": "<|endoftext|>",
|
| 236 |
+
"processor_class": "Qwen3VLProcessor",
|
| 237 |
+
"split_special_tokens": false,
|
| 238 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 239 |
+
"unk_token": null
|
| 240 |
+
}
|
processor/video_preprocessor_config.json
ADDED
|
@@ -0,0 +1,41 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"crop_size": null,
|
| 3 |
+
"data_format": "channels_first",
|
| 4 |
+
"default_to_square": true,
|
| 5 |
+
"device": null,
|
| 6 |
+
"do_center_crop": null,
|
| 7 |
+
"do_convert_rgb": true,
|
| 8 |
+
"do_normalize": true,
|
| 9 |
+
"do_rescale": true,
|
| 10 |
+
"do_resize": true,
|
| 11 |
+
"do_sample_frames": true,
|
| 12 |
+
"fps": 2,
|
| 13 |
+
"image_mean": [
|
| 14 |
+
0.5,
|
| 15 |
+
0.5,
|
| 16 |
+
0.5
|
| 17 |
+
],
|
| 18 |
+
"image_std": [
|
| 19 |
+
0.5,
|
| 20 |
+
0.5,
|
| 21 |
+
0.5
|
| 22 |
+
],
|
| 23 |
+
"input_data_format": null,
|
| 24 |
+
"max_frames": 768,
|
| 25 |
+
"merge_size": 2,
|
| 26 |
+
"min_frames": 4,
|
| 27 |
+
"num_frames": null,
|
| 28 |
+
"pad_size": null,
|
| 29 |
+
"patch_size": 16,
|
| 30 |
+
"processor_class": "Qwen3VLProcessor",
|
| 31 |
+
"resample": 3,
|
| 32 |
+
"rescale_factor": 0.00392156862745098,
|
| 33 |
+
"return_metadata": false,
|
| 34 |
+
"size": {
|
| 35 |
+
"longest_edge": 25165824,
|
| 36 |
+
"shortest_edge": 4096
|
| 37 |
+
},
|
| 38 |
+
"temporal_patch_size": 2,
|
| 39 |
+
"video_metadata": null,
|
| 40 |
+
"video_processor_type": "Qwen3VLVideoProcessor"
|
| 41 |
+
}
|
processor/vocab.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
special_tokens_map.json
ADDED
|
@@ -0,0 +1,31 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"additional_special_tokens": [
|
| 3 |
+
"<|im_start|>",
|
| 4 |
+
"<|im_end|>",
|
| 5 |
+
"<|object_ref_start|>",
|
| 6 |
+
"<|object_ref_end|>",
|
| 7 |
+
"<|box_start|>",
|
| 8 |
+
"<|box_end|>",
|
| 9 |
+
"<|quad_start|>",
|
| 10 |
+
"<|quad_end|>",
|
| 11 |
+
"<|vision_start|>",
|
| 12 |
+
"<|vision_end|>",
|
| 13 |
+
"<|vision_pad|>",
|
| 14 |
+
"<|image_pad|>",
|
| 15 |
+
"<|video_pad|>"
|
| 16 |
+
],
|
| 17 |
+
"eos_token": {
|
| 18 |
+
"content": "<|im_end|>",
|
| 19 |
+
"lstrip": false,
|
| 20 |
+
"normalized": false,
|
| 21 |
+
"rstrip": false,
|
| 22 |
+
"single_word": false
|
| 23 |
+
},
|
| 24 |
+
"pad_token": {
|
| 25 |
+
"content": "<|endoftext|>",
|
| 26 |
+
"lstrip": false,
|
| 27 |
+
"normalized": false,
|
| 28 |
+
"rstrip": false,
|
| 29 |
+
"single_word": false
|
| 30 |
+
}
|
| 31 |
+
}
|
text_encoder/config.json
ADDED
|
@@ -0,0 +1,65 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"Qwen3VLForConditionalGeneration"
|
| 4 |
+
],
|
| 5 |
+
"dtype": "bfloat16",
|
| 6 |
+
"image_token_id": 151655,
|
| 7 |
+
"model_type": "qwen3_vl",
|
| 8 |
+
"text_config": {
|
| 9 |
+
"attention_bias": false,
|
| 10 |
+
"attention_dropout": 0.0,
|
| 11 |
+
"bos_token_id": 151643,
|
| 12 |
+
"dtype": "bfloat16",
|
| 13 |
+
"eos_token_id": 151645,
|
| 14 |
+
"head_dim": 128,
|
| 15 |
+
"hidden_act": "silu",
|
| 16 |
+
"hidden_size": 4096,
|
| 17 |
+
"initializer_range": 0.02,
|
| 18 |
+
"intermediate_size": 12288,
|
| 19 |
+
"max_position_embeddings": 262144,
|
| 20 |
+
"model_type": "qwen3_vl_text",
|
| 21 |
+
"num_attention_heads": 32,
|
| 22 |
+
"num_hidden_layers": 36,
|
| 23 |
+
"num_key_value_heads": 8,
|
| 24 |
+
"pad_token_id": null,
|
| 25 |
+
"rms_norm_eps": 1e-06,
|
| 26 |
+
"rope_parameters": {
|
| 27 |
+
"mrope_interleaved": true,
|
| 28 |
+
"mrope_section": [
|
| 29 |
+
24,
|
| 30 |
+
20,
|
| 31 |
+
20
|
| 32 |
+
],
|
| 33 |
+
"rope_theta": 5000000,
|
| 34 |
+
"rope_type": "default"
|
| 35 |
+
},
|
| 36 |
+
"use_cache": true,
|
| 37 |
+
"vocab_size": 151936
|
| 38 |
+
},
|
| 39 |
+
"tie_word_embeddings": false,
|
| 40 |
+
"transformers_version": "5.3.0.dev0",
|
| 41 |
+
"video_token_id": 151656,
|
| 42 |
+
"vision_config": {
|
| 43 |
+
"deepstack_visual_indexes": [
|
| 44 |
+
8,
|
| 45 |
+
16,
|
| 46 |
+
24
|
| 47 |
+
],
|
| 48 |
+
"depth": 27,
|
| 49 |
+
"dtype": "bfloat16",
|
| 50 |
+
"hidden_act": "gelu_pytorch_tanh",
|
| 51 |
+
"hidden_size": 1152,
|
| 52 |
+
"in_channels": 3,
|
| 53 |
+
"initializer_range": 0.02,
|
| 54 |
+
"intermediate_size": 4304,
|
| 55 |
+
"model_type": "qwen3_vl",
|
| 56 |
+
"num_heads": 16,
|
| 57 |
+
"num_position_embeddings": 2304,
|
| 58 |
+
"out_hidden_size": 4096,
|
| 59 |
+
"patch_size": 16,
|
| 60 |
+
"spatial_merge_size": 2,
|
| 61 |
+
"temporal_patch_size": 2
|
| 62 |
+
},
|
| 63 |
+
"vision_end_token_id": 151653,
|
| 64 |
+
"vision_start_token_id": 151652
|
| 65 |
+
}
|
text_encoder/generation_config.json
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"bos_token_id": 151643,
|
| 3 |
+
"do_sample": true,
|
| 4 |
+
"eos_token_id": [
|
| 5 |
+
151645,
|
| 6 |
+
151643
|
| 7 |
+
],
|
| 8 |
+
"pad_token_id": 151643,
|
| 9 |
+
"temperature": 0.7,
|
| 10 |
+
"top_k": 20,
|
| 11 |
+
"top_p": 0.8,
|
| 12 |
+
"transformers_version": "5.3.0.dev0"
|
| 13 |
+
}
|
text_encoder/model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ba4cadb5ccea2f11dcbb7facce01cb35a76c9a2d5553b34f681b2e16227a1f02
|
| 3 |
+
size 16289680712
|
tokenizer.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:aeb13307a71acd8fe81861d94ad54ab689df773318809eed3cbe794b4492dae4
|
| 3 |
+
size 11422654
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,240 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_bos_token": false,
|
| 3 |
+
"add_prefix_space": false,
|
| 4 |
+
"added_tokens_decoder": {
|
| 5 |
+
"151643": {
|
| 6 |
+
"content": "<|endoftext|>",
|
| 7 |
+
"lstrip": false,
|
| 8 |
+
"normalized": false,
|
| 9 |
+
"rstrip": false,
|
| 10 |
+
"single_word": false,
|
| 11 |
+
"special": true
|
| 12 |
+
},
|
| 13 |
+
"151644": {
|
| 14 |
+
"content": "<|im_start|>",
|
| 15 |
+
"lstrip": false,
|
| 16 |
+
"normalized": false,
|
| 17 |
+
"rstrip": false,
|
| 18 |
+
"single_word": false,
|
| 19 |
+
"special": true
|
| 20 |
+
},
|
| 21 |
+
"151645": {
|
| 22 |
+
"content": "<|im_end|>",
|
| 23 |
+
"lstrip": false,
|
| 24 |
+
"normalized": false,
|
| 25 |
+
"rstrip": false,
|
| 26 |
+
"single_word": false,
|
| 27 |
+
"special": true
|
| 28 |
+
},
|
| 29 |
+
"151646": {
|
| 30 |
+
"content": "<|object_ref_start|>",
|
| 31 |
+
"lstrip": false,
|
| 32 |
+
"normalized": false,
|
| 33 |
+
"rstrip": false,
|
| 34 |
+
"single_word": false,
|
| 35 |
+
"special": true
|
| 36 |
+
},
|
| 37 |
+
"151647": {
|
| 38 |
+
"content": "<|object_ref_end|>",
|
| 39 |
+
"lstrip": false,
|
| 40 |
+
"normalized": false,
|
| 41 |
+
"rstrip": false,
|
| 42 |
+
"single_word": false,
|
| 43 |
+
"special": true
|
| 44 |
+
},
|
| 45 |
+
"151648": {
|
| 46 |
+
"content": "<|box_start|>",
|
| 47 |
+
"lstrip": false,
|
| 48 |
+
"normalized": false,
|
| 49 |
+
"rstrip": false,
|
| 50 |
+
"single_word": false,
|
| 51 |
+
"special": true
|
| 52 |
+
},
|
| 53 |
+
"151649": {
|
| 54 |
+
"content": "<|box_end|>",
|
| 55 |
+
"lstrip": false,
|
| 56 |
+
"normalized": false,
|
| 57 |
+
"rstrip": false,
|
| 58 |
+
"single_word": false,
|
| 59 |
+
"special": true
|
| 60 |
+
},
|
| 61 |
+
"151650": {
|
| 62 |
+
"content": "<|quad_start|>",
|
| 63 |
+
"lstrip": false,
|
| 64 |
+
"normalized": false,
|
| 65 |
+
"rstrip": false,
|
| 66 |
+
"single_word": false,
|
| 67 |
+
"special": true
|
| 68 |
+
},
|
| 69 |
+
"151651": {
|
| 70 |
+
"content": "<|quad_end|>",
|
| 71 |
+
"lstrip": false,
|
| 72 |
+
"normalized": false,
|
| 73 |
+
"rstrip": false,
|
| 74 |
+
"single_word": false,
|
| 75 |
+
"special": true
|
| 76 |
+
},
|
| 77 |
+
"151652": {
|
| 78 |
+
"content": "<|vision_start|>",
|
| 79 |
+
"lstrip": false,
|
| 80 |
+
"normalized": false,
|
| 81 |
+
"rstrip": false,
|
| 82 |
+
"single_word": false,
|
| 83 |
+
"special": true
|
| 84 |
+
},
|
| 85 |
+
"151653": {
|
| 86 |
+
"content": "<|vision_end|>",
|
| 87 |
+
"lstrip": false,
|
| 88 |
+
"normalized": false,
|
| 89 |
+
"rstrip": false,
|
| 90 |
+
"single_word": false,
|
| 91 |
+
"special": true
|
| 92 |
+
},
|
| 93 |
+
"151654": {
|
| 94 |
+
"content": "<|vision_pad|>",
|
| 95 |
+
"lstrip": false,
|
| 96 |
+
"normalized": false,
|
| 97 |
+
"rstrip": false,
|
| 98 |
+
"single_word": false,
|
| 99 |
+
"special": true
|
| 100 |
+
},
|
| 101 |
+
"151655": {
|
| 102 |
+
"content": "<|image_pad|>",
|
| 103 |
+
"lstrip": false,
|
| 104 |
+
"normalized": false,
|
| 105 |
+
"rstrip": false,
|
| 106 |
+
"single_word": false,
|
| 107 |
+
"special": true
|
| 108 |
+
},
|
| 109 |
+
"151656": {
|
| 110 |
+
"content": "<|video_pad|>",
|
| 111 |
+
"lstrip": false,
|
| 112 |
+
"normalized": false,
|
| 113 |
+
"rstrip": false,
|
| 114 |
+
"single_word": false,
|
| 115 |
+
"special": true
|
| 116 |
+
},
|
| 117 |
+
"151657": {
|
| 118 |
+
"content": "<tool_call>",
|
| 119 |
+
"lstrip": false,
|
| 120 |
+
"normalized": false,
|
| 121 |
+
"rstrip": false,
|
| 122 |
+
"single_word": false,
|
| 123 |
+
"special": false
|
| 124 |
+
},
|
| 125 |
+
"151658": {
|
| 126 |
+
"content": "</tool_call>",
|
| 127 |
+
"lstrip": false,
|
| 128 |
+
"normalized": false,
|
| 129 |
+
"rstrip": false,
|
| 130 |
+
"single_word": false,
|
| 131 |
+
"special": false
|
| 132 |
+
},
|
| 133 |
+
"151659": {
|
| 134 |
+
"content": "<|fim_prefix|>",
|
| 135 |
+
"lstrip": false,
|
| 136 |
+
"normalized": false,
|
| 137 |
+
"rstrip": false,
|
| 138 |
+
"single_word": false,
|
| 139 |
+
"special": false
|
| 140 |
+
},
|
| 141 |
+
"151660": {
|
| 142 |
+
"content": "<|fim_middle|>",
|
| 143 |
+
"lstrip": false,
|
| 144 |
+
"normalized": false,
|
| 145 |
+
"rstrip": false,
|
| 146 |
+
"single_word": false,
|
| 147 |
+
"special": false
|
| 148 |
+
},
|
| 149 |
+
"151661": {
|
| 150 |
+
"content": "<|fim_suffix|>",
|
| 151 |
+
"lstrip": false,
|
| 152 |
+
"normalized": false,
|
| 153 |
+
"rstrip": false,
|
| 154 |
+
"single_word": false,
|
| 155 |
+
"special": false
|
| 156 |
+
},
|
| 157 |
+
"151662": {
|
| 158 |
+
"content": "<|fim_pad|>",
|
| 159 |
+
"lstrip": false,
|
| 160 |
+
"normalized": false,
|
| 161 |
+
"rstrip": false,
|
| 162 |
+
"single_word": false,
|
| 163 |
+
"special": false
|
| 164 |
+
},
|
| 165 |
+
"151663": {
|
| 166 |
+
"content": "<|repo_name|>",
|
| 167 |
+
"lstrip": false,
|
| 168 |
+
"normalized": false,
|
| 169 |
+
"rstrip": false,
|
| 170 |
+
"single_word": false,
|
| 171 |
+
"special": false
|
| 172 |
+
},
|
| 173 |
+
"151664": {
|
| 174 |
+
"content": "<|file_sep|>",
|
| 175 |
+
"lstrip": false,
|
| 176 |
+
"normalized": false,
|
| 177 |
+
"rstrip": false,
|
| 178 |
+
"single_word": false,
|
| 179 |
+
"special": false
|
| 180 |
+
},
|
| 181 |
+
"151665": {
|
| 182 |
+
"content": "<tool_response>",
|
| 183 |
+
"lstrip": false,
|
| 184 |
+
"normalized": false,
|
| 185 |
+
"rstrip": false,
|
| 186 |
+
"single_word": false,
|
| 187 |
+
"special": false
|
| 188 |
+
},
|
| 189 |
+
"151666": {
|
| 190 |
+
"content": "</tool_response>",
|
| 191 |
+
"lstrip": false,
|
| 192 |
+
"normalized": false,
|
| 193 |
+
"rstrip": false,
|
| 194 |
+
"single_word": false,
|
| 195 |
+
"special": false
|
| 196 |
+
},
|
| 197 |
+
"151667": {
|
| 198 |
+
"content": "<think>",
|
| 199 |
+
"lstrip": false,
|
| 200 |
+
"normalized": false,
|
| 201 |
+
"rstrip": false,
|
| 202 |
+
"single_word": false,
|
| 203 |
+
"special": false
|
| 204 |
+
},
|
| 205 |
+
"151668": {
|
| 206 |
+
"content": "</think>",
|
| 207 |
+
"lstrip": false,
|
| 208 |
+
"normalized": false,
|
| 209 |
+
"rstrip": false,
|
| 210 |
+
"single_word": false,
|
| 211 |
+
"special": false
|
| 212 |
+
}
|
| 213 |
+
},
|
| 214 |
+
"additional_special_tokens": [
|
| 215 |
+
"<|im_start|>",
|
| 216 |
+
"<|im_end|>",
|
| 217 |
+
"<|object_ref_start|>",
|
| 218 |
+
"<|object_ref_end|>",
|
| 219 |
+
"<|box_start|>",
|
| 220 |
+
"<|box_end|>",
|
| 221 |
+
"<|quad_start|>",
|
| 222 |
+
"<|quad_end|>",
|
| 223 |
+
"<|vision_start|>",
|
| 224 |
+
"<|vision_end|>",
|
| 225 |
+
"<|vision_pad|>",
|
| 226 |
+
"<|image_pad|>",
|
| 227 |
+
"<|video_pad|>"
|
| 228 |
+
],
|
| 229 |
+
"bos_token": null,
|
| 230 |
+
"clean_up_tokenization_spaces": false,
|
| 231 |
+
"eos_token": "<|im_end|>",
|
| 232 |
+
"errors": "replace",
|
| 233 |
+
"extra_special_tokens": {},
|
| 234 |
+
"model_max_length": 262144,
|
| 235 |
+
"pad_token": "<|endoftext|>",
|
| 236 |
+
"processor_class": "Qwen3VLProcessor",
|
| 237 |
+
"split_special_tokens": false,
|
| 238 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 239 |
+
"unk_token": null
|
| 240 |
+
}
|
video_preprocessor_config.json
ADDED
|
@@ -0,0 +1,41 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"crop_size": null,
|
| 3 |
+
"data_format": "channels_first",
|
| 4 |
+
"default_to_square": true,
|
| 5 |
+
"device": null,
|
| 6 |
+
"do_center_crop": null,
|
| 7 |
+
"do_convert_rgb": true,
|
| 8 |
+
"do_normalize": true,
|
| 9 |
+
"do_rescale": true,
|
| 10 |
+
"do_resize": true,
|
| 11 |
+
"do_sample_frames": true,
|
| 12 |
+
"fps": 2,
|
| 13 |
+
"image_mean": [
|
| 14 |
+
0.5,
|
| 15 |
+
0.5,
|
| 16 |
+
0.5
|
| 17 |
+
],
|
| 18 |
+
"image_std": [
|
| 19 |
+
0.5,
|
| 20 |
+
0.5,
|
| 21 |
+
0.5
|
| 22 |
+
],
|
| 23 |
+
"input_data_format": null,
|
| 24 |
+
"max_frames": 768,
|
| 25 |
+
"merge_size": 2,
|
| 26 |
+
"min_frames": 4,
|
| 27 |
+
"num_frames": null,
|
| 28 |
+
"pad_size": null,
|
| 29 |
+
"patch_size": 16,
|
| 30 |
+
"processor_class": "Qwen3VLProcessor",
|
| 31 |
+
"resample": 3,
|
| 32 |
+
"rescale_factor": 0.00392156862745098,
|
| 33 |
+
"return_metadata": false,
|
| 34 |
+
"size": {
|
| 35 |
+
"longest_edge": 25165824,
|
| 36 |
+
"shortest_edge": 4096
|
| 37 |
+
},
|
| 38 |
+
"temporal_patch_size": 2,
|
| 39 |
+
"video_metadata": null,
|
| 40 |
+
"video_processor_type": "Qwen3VLVideoProcessor"
|
| 41 |
+
}
|
vocab.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|