IC-761GPT-3.5-turbo and GPT-4 are susceptible to specific circulating jailbreaking prompts, with 'jailmommy' achieving a 71.16% success rate in producing toxic outputs
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, Yuntian Deng
The paper analysed 1,009,245 real user–chatbot conversations collected from a service powered by GPT-3.5-turbo and GPT-4 APIs. Seven jailbreaking prompts that circulate on social media were identified and their success rates measured by checking whether the chatbot's response was flagged as toxic by either Detoxify or the OpenAI Moderation API. The 'jailmommy' prompt achieved the highest success rate at 71.16%, followed by 'nsfwgpt' at 68.34% and 'eroticachan' at 65.91%. The most frequently used prompt, 'narotica' (3,903 occurrences from 211 users), succeeded 61.82% of the time. Overall, 6.58% of chatbot turns in the dataset were flagged as toxic by at least one detector.
Evidence
correlational
Key metric
jailmommy 71.16%, nsfwgpt 68.34%, eroticachan 65.91%, narotica 61.82%, 4chan user 60.78%, alphabreak 38.42%, do anything now 15.83%; overall chatbot toxicity 6.58% (either detector)
Caveat
Success is defined by flagging from either Detoxify or OpenAI Moderation API, which may not capture all harmful outputs. The analysis is on a combined dataset from both GPT-3.5-turbo and GPT-4 without a per-model breakdown. The paper notes that anonymity of the service may attract users more likely to attempt jailbreaking, introducing selection bias.