IC-363LLM inferences about others' preferences from observed decisions are highly correlated with human inferences because both assume the decision-maker is rational

Ryan Liu, Jiayi Geng, Joshua Peterson, Ilia Sucholutsky, Thomas L. Griffiths

SourceLarge Language Models Assume People are More Rational than We Really are

Using 47 observed decisions from Jern et al. (2017), LLMs were asked to rank decisions by how strongly they suggest a preference for a target item, via pairwise comparisons. LLM inferences correlated positively with rational models (absolute and relative utility), with correlation increasing with model capability and reasoning: Llama-3-8B zero-shot 0.20, Llama-3-8B CoT 0.62, Llama-3-70B CoT 0.88/0.89, GPT-4o CoT 0.95/0.94. Crucially, LLM inferences were even more correlated with human inferences than with rational models: GPT-4o CoT achieved ρ = 0.97 with humans versus 0.95 with absolute utility. This alignment arises because humans also ascribe rational decision-making to others when interpreting their choices, so the LLMs' rational assumption matches human expectation rather than human behavior.

Evidence
correlational
Key metric
GPT-4o CoT: ρ = 0.97 with humans, 0.95 with absolute utility, 0.94 with relative utility; Llama-3-8B zero-shot: 0.20 with absolute utility; Llama-3-70B CoT: 0.88/0.89; humans: 0.98 with absolute utility
Caveat
Only 47 decisions were used, limiting statistical power. Claude 3 Opus used an artificial sample size of 5 due to cost constraints. The negative (electric shock) context showed lower correlations overall, and LLM correlations with humans exceeded those with rational models in that context, suggesting shared heuristics beyond pure rationality.
Model
Llama 3 8B, 70B, Claude 3 Opus, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4 Turbo, GPT-4o
Methods
Chain-of-Thought prompting / Chain-of-Thought (CoT) prompting / CoT prompting / Few-shot Chain of Thought / Few-shot CoT / Chain-of-Thought (CoT-S) / CoT-bag prompting / Few-shot CoT prompting / Rationale prompting / Zero-shot CoT prompting / Wei et al. (2022) chain-of-thought prompting [primary]
Related findings
IC-362
Extraction
automatic-extraction