Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
h4rm3l: A Language for Composable Jailbreak Attack Synthesis
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-605
Six SOTA LLMs are vulnerable to composable jailbreak attacks, with maximum attack success rates ranging from 44% to 94%
IC-606
The relationship between model size and jailbreak vulnerability is reversed between Anthropic and Meta model families