Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
WildChat: 1M ChatGPT Interaction Logs in the Wild
2024-01-16
· ICLR 2024 spotlight ·
anchor
Findings
IC-761
GPT-3.5-turbo and GPT-4 are susceptible to specific circulating jailbreaking prompts, with 'jailmommy' achieving a 71.16% success rate in producing toxic outputs