IC-30756 LLMs on Sorry-Bench show fulfillment rates ranging from below 10% (Claude-2, Gemini-1.5) to above 90% (Mistral-7B-instruct-v0.1, Dolphin-2.6-mixtral-8x7b), with GPT-4o at 30% and Llama-3-70B at 35%

Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, Prateek Mittal

SourceSORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

The paper benchmarks 56 proprietary and open-weight LLMs on 440 class-balanced unsafe instructions across 44 fine-grained safety categories. Fulfillment rates (the fraction of unsafe instructions the model complies with) vary dramatically: Claude-2.1 and Claude-2.0 refuse almost all instructions in the first 24 categories, Gemini-1.5-flash refuses all 'potentially unqualified advice' requests, while Mistral-7B-instruct-v0.1 and Dolphin-2.6-mixtral-8x7b fulfill over 90% of instructions. GPT-4o shows a 30% fulfillment rate, higher than its predecessor GPT-3.5-turbo-1106 at 11%, which the authors attribute to OpenAI's updated model spec allowing more permissive responses on certain topics.

Evidence
correlational
Key metric
fulfillment rates: Claude-2 <10%, Gemini-1.5 <10%, GPT-4o 30%, Llama-3-70B 35%, Mistral-7B-instruct-v0.2 67%, Mistral-7B-instruct-v0.1 >90%, Dolphin-2.6-mixtral-8x7b >90%, Llama-2-70B-chat 12%, Gemini-pro 33%
Caveat
Results depend on evaluation configuration (temperature 0.7, no system prompt for most models); the authors note that varying decoding parameters can noticeably impact model safety behavior. The benchmark does not capture multi-category unsafe scenarios or neutral prompts.
Model
GPT-4o, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo, Claude 2.1, Claude 2.0, Gemini 1.5 / Gemini Pro 1.5 Gemini 1.5 Flash, Gemini Pro, Llama 3 70B Instruct, 8B Instruct, Llama 2 / Llama 2 base Llama-2-70B-Chat, Llama 2 7B Chat / Llama-2-chat-7b, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 Mistral 7B Instruct v0.2, Mistral-7B-Instruct-v0.1, Gemma Gemma-7B-IT, Vicuna Vicuna-7B-v1.5, OpenChat-3.5-0106, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b Dolphin 2.6 Mixtral 8x7B, Zephyr-7B-beta
Datasets
SorryBench [eval]
Methods
Cohen's kappa [eval]
Related work
HarmBench / HarmBench Prompt / HarmBench Response / HarmBench-adv [compared-to], Salad-Bench [compared-to], WildGuard [compared-to]
Related findings
IC-308, IC-309, IC-310
Extraction
automatic-extraction