IC-605Six SOTA LLMs are vulnerable to composable jailbreak attacks, with maximum attack success rates ranging from 44% to 94%

Moussa Koulako Bala Doumbouya, Ananjan Nandi, Gabriel Poesia, Davide Ghilardi, Anna Goldie, Federico Bianchi, Dan Jurafsky, Christopher D Manning

Sourceh4rm3l: A Language for Composable Jailbreak Attack Synthesis

The paper benchmarks GPT-3.5, GPT-4o, Claude-3-Sonnet, Claude-3-Haiku, Llama-3-8B, and Llama-3-70B against 83 jailbreak attacks (identity, 22 state-of-the-art, and 60 synthesized) using 50 illicit prompts from AdvBench. The highest recorded attack success rates are 88% for GPT-3.5, 94% for GPT-4o, 82% for Claude-3-Haiku, 44% for Claude-3-Sonnet, 90% for Llama-3-70B, and 74% for Llama-3-8B. The most effective attacks targeting a given model are rarely as effective against other models, indicating model-specific vulnerability profiles. Synthesized attacks outperform the best state-of-the-art attacks by margins of 10–50 percentage points across all six models.

Evidence
correlational
Key metric
highest recorded ASR: 88% for gpt-3.5, 94% for gpt-4o, 82% for claude-3-haiku, 44% for claude-3-sonnet, 90% for llama-3-70b, and 74% for llama-3-8b; synthesized attacks outperform best SOTA attacks by 10% (gpt-3.5), 50% (gpt-4o), 42% (claude-3-haiku), 42% (claude-3-sonnet), 32% (llama-3-70b), 46% (llama-3-8b)
Caveat
Anthropic applied additional safety filters to the authors' account during experiments, so Claude-3 results reflect behaviour under that extra layer of protection and are not directly comparable to the other models.
Model
GPT-3.5 / ChatGPT-3.5, GPT-4o, Claude 3 Sonnet, Haiku, Llama 3 8B, 70B
Concepts
Failure mode
Datasets
AdvBench / AdvBench-50 [eval]
Related work
HarmBench / HarmBench Prompt / HarmBench Response / HarmBench-adv [compared-to], JailbreakBench [compared-to]
Related findings
IC-606
Extraction
automatic-extraction