The paper evaluates the autoencoder components of previously released Stable Diffusion models on the COCO 2017 validation split at 256x256 pixels. SD-VAE 1.x achieves PSNR 23.4, SSIM 0.69, LPIPS 0.96, rFID 5.0, while SD-VAE 2.x achieves PSNR 24.5, SSIM 0.71, LPIPS 0.92, rFID 4.7. The new SDXL-VAE outperforms both on all four metrics.
The paper notes that SD 2.x uses an improved version of SD 1.x's autoencoder with a reduced perceptual loss weight and more compute; the new VAE is trained from scratch.