IC-308Linguistic mutations to unsafe prompts significantly and inconsistently alter safety refusal across models, with persuasion techniques increasing fulfillment by 5-66% and encoding/encryption decreasing it by 15-68%

Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, Prateek Mittal

SourceSORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

The paper applies 20 linguistic mutations (writing styles, persuasion techniques, encoding/encryption, multi-language translations) to the 440 base unsafe instructions and measures how fulfillment rates change for 8 models. Persuasion techniques (logical appeal, authority endorsement, misrepresentation, evidence-based, expert endorsement) consistently increase fulfillment across all models. Encoding and encryption (ASCII, Caesar, Morse, Atbash) cause most models to output nonsense, reducing fulfillment by 15-68%, though GPT-4o shows increased fulfillment (+11-16%) for 2 of 4 strategies. Multi-language effects are model-dependent: GPT-4o maintains <±4% variation across languages, while Vicuna, Mistral, and OpenChat show marked decreases (-20 to -58%) for low-resource languages.

Evidence
correlational
Key metric
persuasion: +5% to +66% fulfillment increase; encoding/encryption: -15% to -68% fulfillment decrease; technical terms: +8% to +19%; GPT-4o multi-language variation <±4%; Mistral-7B-instruct-v0.2 Atbash: -0.67; OpenChat-3.5-0106 Atbash: -0.68
Caveat
Only 8 models were evaluated over the 20 linguistic mutations due to computational constraints. The multi-language translations were generated via Google Translate API, which may introduce translation artifacts.
Model
GPT-4o, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo, Llama 3 8B Instruct, 70B Instruct, Gemma Gemma-7B-IT, Vicuna Vicuna-7B-v1.5, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 Mistral 7B Instruct v0.2, OpenChat-3.5-0106
Concepts
Failure mode
Datasets
SorryBench [eval]
Methods
Cohen's kappa [eval]
Related work
Rainbow Teaming [context], How Johnny Can Persuade LLMs to Jailbreak Them [context]
Related findings
IC-307, IC-309, IC-310
Extraction
automatic-extraction