IC-828Jailbreaking GPT-4 via cipherchat shifts its psychological profile toward human norms and reduces emotional intelligence scores

Jen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho LAM, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, Michael Lyu

SourceOn the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMs

Applying a Caesar cipher shift-3 jailbreak (cipherchat) to GPT-4's prompts bypasses its safety alignment and produces a different psychological profile (GPT-4-JB). The jailbroken model's BFI openness (3.8±0.6) moves closer to the human norm (3.9±0.7) compared to default GPT-4 (4.2±0.6). The jailbreak causes a substantial reduction in EIS (from 151.4±18.7 to 121.8±12.0) and empathy (from 6.8±0.4 to 4.6±0.2), though WLEIS subscales show no statistically significant change. The authors interpret this as revealing GPT-4's 'intrinsic' psychological nature beneath safety constraints.

Evidence
correlational
Key metric
EIS: GPT-4 151.4±18.7 vs GPT-4-JB 121.8±12.0; Empathy: GPT-4 6.8±0.4 vs GPT-4-JB 4.6±0.2; BFI openness: GPT-4 4.2±0.6 vs GPT-4-JB 3.8±0.6 vs human 3.9±0.7
Caveat
The jailbreak is a prompting technique, not a modification of model weights; the observed shift may reflect the model's response to an unusual input format rather than a removal of a constraint. Only one cipher variant (Caesar shift 3) was tested.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
Methods
CipherChat [primary]
Related findings
IC-827, IC-829
Extraction
automatic-extraction