IC-497Text-to-image models experience performance drops exceeding 10% under adversarial prompts, with spatial reasoning being the most vulnerable task across all models

Chejian Xu, Jiawei Zhang, Zhaorun Chen, Chulin Xie, Mintong Kang, Yujin Potter, Zhun Wang, Zhuowen Yuan, Alexander Xiong, Zidi Xiong, Chenhui Zhang, Lingzhi Yuan, Yi Zeng, Peiyang Xu, Chengquan Guo, Andy Zhou, Jeffrey Ziwei Tan, Xuandong Zhao, Francesco Pinto, Zhen Xiang, Yu Gai, Zinan Lin, Dan Hendrycks, Bo Li, Dawn Song

SourceMMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models

The paper constructs adversarial prompts using GCg and MMP attacks against surrogate models, then evaluates transferability to target text-to-image and image-to-text models. All text-to-image models show performance drops above 10% from their benign accuracy. The best text-to-image model (Nova Canvas) achieves only 62.42% averaged robust accuracy. For image-to-text models, spatial reasoning is the most challenging task, where the best model (GPT-4o) achieves only 53.79% accuracy. Newer models within the same family (DALL·E 3 vs DALL·E 2, GPT-4o vs GPT-4V) show both higher benign accuracy and better robustness.

Evidence
correlational
Key metric
best t2i model nova canvas achieves 62.42% averaged robust accuracy; best i2t model gpt-4o achieves 53.79% on spatial reasoning; dall·e 3 robust accuracy 61.38% vs dall·e 2 46.66%; gpt-4o robust accuracy 90.04% vs gpt-4v 85.27%
Caveat
Adversarial examples were generated against surrogate models and transferred to target models, so the attack strength may not reflect optimal white-box attacks against each target.
Model
Nova Canvas, FLUX / FLUX1, DALL·E 3, DALL·E 2, GPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4V / GPT-4 vision, Stable Diffusion SDXL, LLaVA-NeXT / LLaVA 1.6
Concepts
Failure mode
Datasets
MS COCO / COCO / COCO 2014 / COCO 2017 / COCO 20k / COCO-it / COCO-wl [source]
Methods
GCG [primary], MMP [primary], AttackVLM [primary]
Related findings
IC-495, IC-496, IC-498
Extraction
automatic-extraction