Diffusers
Safetensors
vae
qwen-image
File size: 4,340 Bytes
98dbe2f
 
 
 
8b6801b
 
 
 
 
98dbe2f
2ad703e
 
 
9e3f392
2ad703e
9e3f392
72371a1
d32449b
18b5e82
d32449b
 
 
 
 
 
8b6801b
9e3f392
 
 
 
 
8b6801b
2ad703e
 
d32449b
 
 
 
8b6801b
 
 
 
 
 
 
 
 
 
 
 
c408d83
8b6801b
c408d83
 
 
8b6801b
c408d83
 
 
8b6801b
 
2ad703e
c408d83
2ad703e
8b6801b
 
 
 
 
 
 
2ad703e
 
 
8b6801b
2ad703e
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
---
license: other
license_name: qwen-research
license_link: https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE
base_model: Qwen/Qwen-Image-2.1
library_name: diffusers
tags:
- vae
- qwen-image
---

# Qwen-Image-2.1-VAE-Texture-Fix

Qwen-Image-2.1-VAE-Texture-Fix is the [Qwen-Image-2.1 VAE](https://huggingface.co/Qwen/Qwen-Image-2.1/tree/main/vae), but finetuned to produce cleaner textures with no checkerboard artifacts.

Qwen-Image-2.1-VAE-Texture-Fix's improved decoding is most noticeable in detailed, photo-style images.

<details>
<summary>Comparison Settings</summary>

The latents for the VAE comparison image below were generated by Qwen-Image-2.1 from the prompt:

> Landscape photograph of a subalpine wildflower meadow in the Pacific Northwest in midsummer: a clear mountain stream winding over mossy boulders through purple lupine and red paintbrush, dense old-growth Douglas fir and western red cedar forest behind, a snow-capped volcano in the distance, golden late-afternoon light, highly detailed

</details>

| Qwen-Image-2.1-VAE ([full-res](./images/decode-qwen-image-2.1-vae-full.png)) | 🪄 Qwen-Image-2.1-VAE-Texture-Fix 🪄 ([full-res](./images/decode-qwen-image-2.1-vae-texture-fix-full.png)) |
| --- | --- |
| ![](./images/decode-qwen-image-2.1-vae-zoomed-1.png) | ![](./images/decode-qwen-image-2.1-vae-texture-fix-zoomed-1.png) |
| ![](./images/decode-qwen-image-2.1-vae-zoomed-2.png) | ![](./images/decode-qwen-image-2.1-vae-texture-fix-zoomed-2.png) |
| ![](./images/decode-qwen-image-2.1-vae.png) | ![](./images/decode-qwen-image-2.1-vae-texture-fix.png) |

# Usage

**ComfyUI**

Download [`qwen_image_2.1_vae_texture_fix_bf16.safetensors`](./qwen_image_2.1_vae_texture_fix_bf16.safetensors) into `ComfyUI/models/vae/` and select it in the `Load VAE` node, in place of `qwen_image_2.1_vae_bf16.safetensors`.

**🧨 Diffusers**

```python
import torch
from diffusers import QwenImage21Pipeline, AutoencoderKLQwenImage21

vae = AutoencoderKLQwenImage21.from_pretrained("madebyollin/qwen-image-2.1-vae-texture-fix", torch_dtype=torch.bfloat16)
pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", vae=vae, torch_dtype=torch.bfloat16).to("cuda")
```

# Mechanism

Qwen-Image-2.1-VAE-Texture-Fix was created by finetuning the Qwen-Image-2.1 VAE decoder for ~5000 steps at learning rate 3e-5, with only the two highest-resolution decoder stages and output head unfrozen (7.5M trainable parameters), using the recipe developed for [TAESD](https://github.com/madebyollin/taesd).

The TAESD recipe, like most image autoencoder training recipes, uses a mix of PSNR-focused (MSE/MAE), LPIPS, and adversarial (GAN) loss terms.
Whenever precise details can't be reconstructed, MSE/MAE loss encourages blurring, LPIPS loss encourages blurring+checkerboarding (among other artifacts), and adversarial loss encourages generating sharp/plausible (but fake) detail without obvious artifacts.
This figure from DC-AE (https://arxiv.org/abs/2410.10733) shows the importance of including adversarial (GAN) loss:

![Demo of the effects of adversarial loss, courtesy of the DC-AE paper](https://cdn-uploads.huggingface.co/production/uploads/630447d40547362a22a969a2/uj9vJSR34LyZEkQeL-VOU.png)

I suspect the original Qwen-Image-2.1-VAE was trained without a working adversarial loss term.

# Metrics

Qwen-Image-2.1-VAE-Texture-Fix makes perceptual quality metrics (rFID) better and reconstruction accuracy metrics (LPIPS/PSNR) slightly worse.

| Metric | Qwen-Image-2.1-VAE | Qwen-Image-2.1-VAE-Texture-Fix |
| --- | --- | --- |
| rFID ↓ (COCO val2017, 5000 images @ 256²) | 3.37 | **2.08** |
| PSNR ↑ (COCO val2017 @ 256²) | **33.30** | 32.86 |
| LPIPS ↓ (COCO val2017 @ 256²) | **0.0357** | 0.0373 |
| PSNR ↑ (DIV2K valid, native 1024² crops) | **32.86** | 32.46 |
| LPIPS ↓ (DIV2K valid, native 1024² crops) | **0.0460** | 0.0480 |

# Attribution Notice

This fine-tuned VAE is based on Qwen/Qwen-Image-2.1; original materials © 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd., licensed under the Qwen RESEARCH LICENSE AGREEMENT (see [LICENSE](./LICENSE)) for non-commercial/research use only. Built with Qwen<sup>*</sup>.

<sup>* In the sense that the **initial VAE weights** are from Qwen-Image. The decoder fine-tuning work was performed by `madebyollin` and Claude Opus.</sup>