Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-927
Human ciphers (ASCII, Unicode, Caesar, Morse) bypass the safety alignment of GPT-4 and GPT-3.5-turbo, with more powerful models producing more unsafe responses
IC-928
SelfCipher (a role-play prompt without explicit cipher rules) evokes a 'secret cipher' in LLMs, achieving high unsafety rates that outperform most human ciphers
IC-929
Simulated character-level ciphers that never appear in pretraining data cannot bypass safety alignment even with 10+ demonstrations