In a user study where participants chose their favourite generation among four models (SDXL with refiner, SDXL base, SD 1.5, SD 2.1), SD 1.5 and SD 2.1 were the least preferred, receiving 7.91% and 6.71% win rates respectively, compared to 48.44% for SDXL with refinement and 36.93% for SDXL base. The paper notes that classical metrics like FID and CLIP scores do not reflect this gap, aligning with findings by Kirstain et al. (2023).
The paper notes that classical metrics (FID, CLIP scores) do not reflect the user-preference gap, suggesting the win rates may be sensitive to specific evaluation conditions.