Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
2025-01-22 · ICLR 2025 Poster · anchor
Findings
- IC-407Safety-aligned LLMs (Llama-2-chat, Llama-3-instruct, Gemma, GPT-3.5, GPT-4o, R2D2) achieve 100% jailbreak attack success rate under adaptive prompt-and-suffix attacks on 50 harmful requests
- IC-408Claude models (2.0, 2.1, 3 Haiku, 3 Sonnet, 3 Opus, 3.5 Sonnet) achieve 100% jailbreak attack success rate under prefilling attacks via the Anthropic API