--- license: other license_name: qwen-research license_link: https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE base_model: Qwen/Qwen-Image-2.1 library_name: diffusers tags: - vae - qwen-image --- # Qwen-Image-2.1-VAE-Texture-Fix Qwen-Image-2.1-VAE-Texture-Fix is the [Qwen-Image-2.1 VAE](https://huggingface.co/Qwen/Qwen-Image-2.1/tree/main/vae), but finetuned to produce cleaner textures with no checkerboard artifacts. Qwen-Image-2.1-VAE-Texture-Fix's improved decoding is most noticeable in detailed, photo-style images.
Comparison Settings The latents for the VAE comparison image below were generated by Qwen-Image-2.1 from the prompt: > Landscape photograph of a subalpine wildflower meadow in the Pacific Northwest in midsummer: a clear mountain stream winding over mossy boulders through purple lupine and red paintbrush, dense old-growth Douglas fir and western red cedar forest behind, a snow-capped volcano in the distance, golden late-afternoon light, highly detailed
| Qwen-Image-2.1-VAE ([full-res](./images/decode-qwen-image-2.1-vae-full.png)) | 🪄 Qwen-Image-2.1-VAE-Texture-Fix 🪄 ([full-res](./images/decode-qwen-image-2.1-vae-texture-fix-full.png)) | | --- | --- | | ![](./images/decode-qwen-image-2.1-vae-zoomed-1.png) | ![](./images/decode-qwen-image-2.1-vae-texture-fix-zoomed-1.png) | | ![](./images/decode-qwen-image-2.1-vae-zoomed-2.png) | ![](./images/decode-qwen-image-2.1-vae-texture-fix-zoomed-2.png) | | ![](./images/decode-qwen-image-2.1-vae.png) | ![](./images/decode-qwen-image-2.1-vae-texture-fix.png) | # Usage **ComfyUI** Download [`qwen_image_2.1_vae_texture_fix_bf16.safetensors`](./qwen_image_2.1_vae_texture_fix_bf16.safetensors) into `ComfyUI/models/vae/` and select it in the `Load VAE` node, in place of `qwen_image_2.1_vae_bf16.safetensors`. **🧨 Diffusers** ```python import torch from diffusers import QwenImage21Pipeline, AutoencoderKLQwenImage21 vae = AutoencoderKLQwenImage21.from_pretrained("madebyollin/qwen-image-2.1-vae-texture-fix", torch_dtype=torch.bfloat16) pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", vae=vae, torch_dtype=torch.bfloat16).to("cuda") ``` # Mechanism Qwen-Image-2.1-VAE-Texture-Fix was created by finetuning the Qwen-Image-2.1 VAE decoder for ~5000 steps at learning rate 3e-5, with only the two highest-resolution decoder stages and output head unfrozen (7.5M trainable parameters), using the recipe developed for [TAESD](https://github.com/madebyollin/taesd). The TAESD recipe, like most image autoencoder training recipes, uses a mix of PSNR-focused (MSE/MAE), LPIPS, and adversarial (GAN) loss terms. Whenever precise details can't be reconstructed, MSE/MAE loss encourages blurring, LPIPS loss encourages blurring+checkerboarding (among other artifacts), and adversarial loss encourages generating sharp/plausible (but fake) detail without obvious artifacts. This figure from DC-AE (https://arxiv.org/abs/2410.10733) shows the importance of including adversarial (GAN) loss: ![Demo of the effects of adversarial loss, courtesy of the DC-AE paper](https://cdn-uploads.huggingface.co/production/uploads/630447d40547362a22a969a2/uj9vJSR34LyZEkQeL-VOU.png) I suspect the original Qwen-Image-2.1-VAE was trained without a working adversarial loss term. # Metrics Qwen-Image-2.1-VAE-Texture-Fix makes perceptual quality metrics (rFID) better and reconstruction accuracy metrics (LPIPS/PSNR) slightly worse. | Metric | Qwen-Image-2.1-VAE | Qwen-Image-2.1-VAE-Texture-Fix | | --- | --- | --- | | rFID ↓ (COCO val2017, 5000 images @ 256²) | 3.37 | **2.08** | | PSNR ↑ (COCO val2017 @ 256²) | **33.30** | 32.86 | | LPIPS ↓ (COCO val2017 @ 256²) | **0.0357** | 0.0373 | | PSNR ↑ (DIV2K valid, native 1024² crops) | **32.86** | 32.46 | | LPIPS ↓ (DIV2K valid, native 1024² crops) | **0.0460** | 0.0480 | # Attribution Notice This fine-tuned VAE is based on Qwen/Qwen-Image-2.1; original materials © 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd., licensed under the Qwen RESEARCH LICENSE AGREEMENT (see [LICENSE](./LICENSE)) for non-commercial/research use only. Built with Qwen*. * In the sense that the **initial VAE weights** are from Qwen-Image. The decoder fine-tuning work was performed by `madebyollin` and Claude Opus.