IC-1204SD-VAE 1.x and SD-VAE 2.x achieve lower reconstruction quality than the new SDXL-VAE on COCO 2017

Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, Robin Rombach

SourceSDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

The paper evaluates the autoencoder components of previously released Stable Diffusion models on the COCO 2017 validation split at 256x256 pixels. SD-VAE 1.x achieves PSNR 23.4, SSIM 0.69, LPIPS 0.96, rFID 5.0, while SD-VAE 2.x achieves PSNR 24.5, SSIM 0.71, LPIPS 0.92, rFID 4.7. The new SDXL-VAE outperforms both on all four metrics.

Evidence
correlational
Key metric
PSNR: sd-vae 1.x 23.4, sd-vae 2.x 24.5; SSIM: 0.69, 0.71; LPIPS: 0.96, 0.92; rFID: 5.0, 4.7 (COCO 2017 val, 256x256)
Caveat
The paper notes that SD 2.x uses an improved version of SD 1.x's autoencoder with a reduced perceptual loss weight and more compute; the new VAE is trained from scratch.
Model
Stable Diffusion 1.5, 2.1
Datasets
MS COCO / COCO / COCO 2014 / COCO 2017 / COCO 20k / COCO-it / COCO-wl [eval]
Related findings
IC-1202, IC-1203
Extraction
automatic-extraction