IC-30756 LLMs on Sorry-Bench show fulfillment rates ranging from below 10% (Claude-2, Gemini-1.5) to above 90% (Mistral-7B-instruct-v0.1, Dolphin-2.6-mixtral-8x7b), with GPT-4o at 30% and Llama-3-70B at 35%
Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, Prateek Mittal
The paper benchmarks 56 proprietary and open-weight LLMs on 440 class-balanced unsafe instructions across 44 fine-grained safety categories. Fulfillment rates (the fraction of unsafe instructions the model complies with) vary dramatically: Claude-2.1 and Claude-2.0 refuse almost all instructions in the first 24 categories, Gemini-1.5-flash refuses all 'potentially unqualified advice' requests, while Mistral-7B-instruct-v0.1 and Dolphin-2.6-mixtral-8x7b fulfill over 90% of instructions. GPT-4o shows a 30% fulfillment rate, higher than its predecessor GPT-3.5-turbo-1106 at 11%, which the authors attribute to OpenAI's updated model spec allowing more permissive responses on certain topics.
Results depend on evaluation configuration (temperature 0.7, no system prompt for most models); the authors note that varying decoding parameters can noticeably impact model safety behavior. The benchmark does not capture multi-category unsafe scenarios or neutral prompts.