Diffusers
Safetensors
vae
qwen-image
madebyollin commited on
Commit
c408d83
·
verified ·
1 Parent(s): 18b5e82

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +8 -4
README.md CHANGED
@@ -48,15 +48,19 @@ pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", vae=vae, torch
48
 
49
  # Mechanism
50
 
51
- Qwen-Image-2.1-VAE-Texture-Fix was created by finetuning the Qwen-Image-2.1 VAE decoder for ~5000 steps at learning rate 3e-5, with only the two highest-resolution decoder stages and output head unfrozen (7.5M trainable parameters), using the recipe developed for [TAESD](https://github.com/madebyollin/taesd). This recipe includes an adversarial loss term which strongly penalizes checkerboarding (and other obvious texture artifacts).
52
 
53
- VAEs are usually trained with a mix of PSNR-focused (MSE/MAE), LPIPS, and adversarial loss terms. Whenever precise details can't be reconstructed, MSE/MAE loss encourages blurring, LPIPS loss encourages blurring+checkerboarding (and various other artifacts), and adversarial loss encourages generating sharp/plausible (but fake) detail without obvious artifacts.
 
 
54
 
55
- I suspect the original Qwen-Image-2.1-VAE was trained without adversarial loss (or with a very weak/nonfunctioning adversarial loss).
 
 
56
 
57
  # Metrics
58
 
59
- Qwen-Image-2.1-VAE-Texture-Fix makes perceptual quality metrics (rFID) better and reconstruction accuracy metrics (LPIPS/PSNR) worse.
60
 
61
  | Metric | Qwen-Image-2.1-VAE | Qwen-Image-2.1-VAE-Texture-Fix |
62
  | --- | --- | --- |
 
48
 
49
  # Mechanism
50
 
51
+ Qwen-Image-2.1-VAE-Texture-Fix was created by finetuning the Qwen-Image-2.1 VAE decoder for ~5000 steps at learning rate 3e-5, with only the two highest-resolution decoder stages and output head unfrozen (7.5M trainable parameters), using the recipe developed for [TAESD](https://github.com/madebyollin/taesd).
52
 
53
+ The TAESD recipe, like most image autoencoder training recipes, uses a mix of PSNR-focused (MSE/MAE), LPIPS, and adversarial (GAN) loss terms.
54
+ Whenever precise details can't be reconstructed, MSE/MAE loss encourages blurring, LPIPS loss encourages blurring+checkerboarding (among other artifacts), and adversarial loss encourages generating sharp/plausible (but fake) detail without obvious artifacts.
55
+ This figure from DC-AE (https://arxiv.org/abs/2410.10733) shows the importance of including adversarial (GAN) loss:
56
 
57
+ ![Demo of the effects of adversarial loss, courtesy of the DC-AE paper](https://cdn-uploads.huggingface.co/production/uploads/630447d40547362a22a969a2/uj9vJSR34LyZEkQeL-VOU.png)
58
+
59
+ I suspect the original Qwen-Image-2.1-VAE was trained without a working adversarial loss term.
60
 
61
  # Metrics
62
 
63
+ Qwen-Image-2.1-VAE-Texture-Fix makes perceptual quality metrics (rFID) better and reconstruction accuracy metrics (LPIPS/PSNR) slightly worse.
64
 
65
  | Metric | Qwen-Image-2.1-VAE | Qwen-Image-2.1-VAE-Texture-Fix |
66
  | --- | --- | --- |