Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Safety Assessment of Chinese Large Language Models
anchor
Findings
IC-927
Human ciphers (ASCII, Unicode, Caesar, Morse) bypass the safety alignment of GPT-4 and GPT-3.5-turbo, with more powerful models producing more unsafe responses
[eval]
IC-928
SelfCipher (a role-play prompt without explicit cipher rules) evokes a 'secret cipher' in LLMs, achieving high unsafety rates that outperform most human ciphers
[eval]
IC-929
Simulated character-level ciphers that never appear in pretraining data cannot bypass safety alignment even with 10+ demonstrations
[eval]