Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
2024-01-16
· ICLR 2024 spotlight ·
anchor
Findings
IC-1461
Varying decoding hyperparameters and removing the system prompt breaks the safety alignment of 9 out of 11 open-source LLMs, raising attack success rate from 0% to over 95%
IC-1462
GPT-3.5-turbo is substantially more robust to the generation exploitation attack, with attack success rate of only 7% compared to over 95% for open-source models