IC-1139Stable Diffusion encodes social biases in its internal concept representations that are not always visually apparent

Hila Chefer, Oran Lang, Mor Geva, Volodymyr Polosukhin, Assaf Shocher, michal Irani, Inbar Mosseri, Lior Wolf

SourceThe Hidden Language of Diffusion Models

Conceptor reveals that SD's internal representations of profession and role concepts include biased demographic associations: 'secretary' decomposes into 'womens, girl, ladies'; 'opera singer' into 'obese, overweight, fat'; 'pastor' into 'nigerian, gospel'; 'journalist' into 'refugee, jews'; 'drinking' into 'millennials, blonde'; 'professor' into 'men'; 'nurse' into 'woman, women, wife, mother'. The authors note these biases are not necessarily observable by simply looking at generated images. They demonstrate that reducing the coefficients of biased tokens while preserving other features can produce debiased generations on the same seeds.

Evidence
correlational
Caveat
The bias examples are illustrative (Table 4 lists 6 concepts); the full extent of biases across all 188 concepts is not systematically quantified. The debiasing demonstration uses only 8 random seeds.
Model
Stable Diffusion
Concepts
Shortcut
Datasets
Bias in Bios [eval]
Methods
Conceptor [primary]
Related work
Luccioni et al. 2023 (Stable Bias) [context]
Related findings
IC-1137, IC-1138, IC-1140
Extraction
automatic-extraction