IC-1342Assigning socio-demographic personas to LLMs causes significant reasoning performance degradation across all four models studied, manifesting as both explicit abstentions and implicit reasoning errors

Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, Tushar Khot

SourceBias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs

The paper assigns 19 socio-demographic personas (spanning race, gender, religion, disability, and political affiliation) to four LLMs via system prompts and measures their accuracy on 24 reasoning datasets. For ChatGPT-3.5 (June 2023), 80% of personas show a statistically significant performance drop relative to a 'human' persona baseline, with drops reaching 64% (physically disabled on high school world history) and 69% (religious on college chemistry). The bias manifests in two ways: explicit abstentions citing stereotypical assumptions (58% of errors for the physically-disabled persona) and implicit reasoning errors on questions where the model does not abstain. The effect is present in all four models but varies in extent: GPT-4-turbo shows bias in 42% of personas, Llama-2-70B-Chat and ChatGPT-3.5 (June) in 80%, and ChatGPT-3.5 (November) in 100%.

Evidence
correlational
Key metric
ChatGPT-3.5 (June): 80% of personas show bias; 64% drop (phys. disabled, high school world history); 69% drop (religious, college chemistry); 33% avg. drop (phys. disabled); 58% of phys. disabled errors are abstentions. GPT-4-turbo: 42% of personas affected. Llama-2-70B-Chat: 80% of personas affected. ChatGPT-3.5 (Nov): 100% of personas affected.
Caveat
The study uses English-language prompts and datasets only. The persona set is not exhaustive and shows a WEIRD preference. Results are averaged across 3 persona instructions and 3 runs, but one instruction exhibits significantly higher bias than the instruction-averaged results.
Model
GPT-3.5 / ChatGPT-3.5, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4 Turbo, Llama 2 / Llama 2 base Llama-2-70B-Chat
Concepts
Failure mode
Datasets
MMLU / MMLU-Math [eval], Big-Bench Hard [eval], MBPP [eval]
Methods
Wilson's confidence interval [eval]
Related findings
IC-1343
Extraction
automatic-extraction