Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-570
Gradient-based image jailbreaks optimized against single or ensemble VLMs are universal for the attacked model(s) but do not transfer to other VLMs, except between highly similar models