Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models
2024-01-16
· ICLR 2024 spotlight ·
anchor
Findings
IC-1438
LLaVA and Llama-Adapter V2 are jailbroken by compositional adversarial images targeting image-based embedding triggers, with near-zero success for textual triggers
IC-1439
LLaVA follows text instructions embedded in adversarial images as if they were user prompts, enabling hidden prompt injection