IC-1203Stable Diffusion 1.5 and 2.1 receive low user-preference win rates (7.91% and 6.71%) in a four-way comparison against SDXL

Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, Robin Rombach

SourceSDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

In a user study where participants chose their favourite generation among four models (SDXL with refiner, SDXL base, SD 1.5, SD 2.1), SD 1.5 and SD 2.1 were the least preferred, receiving 7.91% and 6.71% win rates respectively, compared to 48.44% for SDXL with refinement and 36.93% for SDXL base. The paper notes that classical metrics like FID and CLIP scores do not reflect this gap, aligning with findings by Kirstain et al. (2023).

Evidence
correlational
Key metric
win rates: sdxl w/ refinement: 48.44%, sdxl base: 36.93%, stable diffusion 1.5: 7.91%, stable diffusion 2.1: 6.71%
Caveat
The paper notes that classical metrics (FID, CLIP scores) do not reflect the user-preference gap, suggesting the win rates may be sensitive to specific evaluation conditions.
Model
Stable Diffusion 1.5, 2.1
Related work
Kirstain et al. 2023 (Pick-a-Pic) [context]
Related findings
IC-1202, IC-1204
Extraction
automatic-extraction