IC-498Multimodal foundation models exhibit severe group unfairness, with race and age biases more pronounced than gender bias in text-to-image models while gender bias is stronger in image-to-text models

Chejian Xu, Jiawei Zhang, Zhaorun Chen, Chulin Xie, Mintong Kang, Yujin Potter, Zhun Wang, Zhuowen Yuan, Alexander Xiong, Zidi Xiong, Chenhui Zhang, Lingzhi Yuan, Yi Zeng, Peiyang Xu, Chengquan Guo, Andy Zhou, Jeffrey Ziwei Tan, Xuandong Zhao, Francesco Pinto, Zhen Xiang, Yu Gai, Zinan Lin, Dan Hendrycks, Bo Li, Dawn Song

SourceMMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models

The paper evaluates fairness across social stereotypes and decision-making scenarios using group unfairness, individual unfairness, and overkill fairness metrics. All text-to-image models show group unfairness scores (gs) between 0.337 and 0.597, far from the ideal of 0. Race and age biases are more pronounced than gender bias in T2I models, while the reverse holds for I2T models. Group unfairness does not observably correlate with individual unfairness. All T2I models also show poor overkill fairness, sacrificing historical accuracy in pursuit of fairness.

Evidence
correlational
Key metric
t2i group unfairness gs: sdxl 0.337, flux 0.597, dall·e 3 0.376; i2t group unfairness gs: gpt-4o 0.142, gpt-4v 0.179, gemini pro-1.5 0.183, llama-3.2 0.033; t2i overkill fairness o: all models between 0.449 and 0.636
Caveat
Llama-3.2's low group unfairness score is partly due to over-refusal (85.2% refusal rate on fairness tasks), not genuine fairness.
Model
FLUX / FLUX1, DALL·E 3, DALL·E 2, Stable Diffusion SDXL, GPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4V / GPT-4 vision, Gemini 1.5 / Gemini Pro 1.5, Llama-3-2-Vision Llama-3.2-90B-Vision-Instruct, Nova Canvas
Concepts
Shortcut, Failure mode
Datasets
UTKFace [source]
Methods
FairFace [eval], CLIPScore [eval]
Related findings
IC-495, IC-496, IC-497
Extraction
automatic-extraction