IC-024Reward model trained on local scaling of Stable Diffusion can guide generation to increase diversity and aesthetic scores

Ahmed Imtiaz Humayun, Ibtihel Amara, Cristina Nader Vasconcelos, Deepak Ramachandran, Candice Schumann, Junfeng He, Katherine A Heller, Golnoosh Farnadi, Negar Rostamzadeh, Mohammad Havaei

SourceWhat Secrets Do Your Manifolds Hold? Understanding the Local Geometry of Generative Models

The paper trains a reward model to predict local scaling level sets for Stable Diffusion. Using this reward model with universal guidance during the reverse diffusion process, the authors show they can increase or decrease the local scaling of generated images. Maximizing the reward leads to added texture, sharpness, contrast, and higher diversity (Vendi score) and human preference scores, while minimizing it leads to blurring and loss of detail. The reward model is trained on 50k to 800k samples from ImageNet latents.

Evidence
interventional
Key metric
Change in Vendi score, change in average local scaling, change in RAHF aesthetic score for increasing pre-guidance local scaling level sets; results shown for reward models trained on 50k, 200k, 400k, 800k samples; guidance strength ρ values (1, 2, -1, -3, -5) used to control effect
Caveat
Training a reward model on local scaling requires computing descriptors for a large dataset; the approach is demonstrated only for Stable Diffusion; the human preference model (RAHF) is trained on synthetic images and may be less reliable for real images.
Model
Stable Diffusion
Datasets
ImageNet-1k / ImageNet / ImageNet-1k-val / ImageNet-Val [train], ImageWoof [eval]
Methods
Universal guidance [primary], Classifier-free guidance / Classifier-free diffusion guidance [compared-to], RAHF (Rich Human Feedback) score [eval], Vendi score [eval]
Extraction
automatic-extraction