The paper benchmarks GPT-3.5, GPT-4o, Claude-3-Sonnet, Claude-3-Haiku, Llama-3-8B, and Llama-3-70B against 83 jailbreak attacks (identity, 22 state-of-the-art, and 60 synthesized) using 50 illicit prompts from AdvBench. The highest recorded attack success rates are 88% for GPT-3.5, 94% for GPT-4o, 82% for Claude-3-Haiku, 44% for Claude-3-Sonnet, 90% for Llama-3-70B, and 74% for Llama-3-8B. The most effective attacks targeting a given model are rarely as effective against other models, indicating model-specific vulnerability profiles. Synthesized attacks outperform the best state-of-the-art attacks by margins of 10–50 percentage points across all six models.
Evidence
correlational
Key metric
highest recorded ASR: 88% for gpt-3.5, 94% for gpt-4o, 82% for claude-3-haiku, 44% for claude-3-sonnet, 90% for llama-3-70b, and 74% for llama-3-8b; synthesized attacks outperform best SOTA attacks by 10% (gpt-3.5), 50% (gpt-4o), 42% (claude-3-haiku), 42% (claude-3-sonnet), 32% (llama-3-70b), 46% (llama-3-8b)
Caveat
Anthropic applied additional safety filters to the authors' account during experiments, so Claude-3 results reflect behaviour under that extra layer of protection and are not directly comparable to the other models.