IC-1267LLMs' alignment with human privacy judgments drops sharply as contextual complexity increases from tier 1 to tier 3

Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, Yejin Choi

SourceCan LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory

The paper measures Pearson correlation between model and human privacy judgments across four tiers of increasing contextual complexity. In tier 1 (simple sensitivity rating), GPT-4 achieves 0.86 and ChatGPT 0.92 correlation. By tier 3 (theory-of-mind scenarios with three parties), GPT-4 drops to 0.10 and ChatGPT to 0.05. All six evaluated models show the same downward trend, with open-source models (Llama-2, Mixtral) starting lower and ending near zero. The authors attribute this to the models' inability to reason about information flow in complex social contexts.

Evidence
correlational
Key metric
GPT-4 Pearson correlation: 0.86 (tier 1), 0.47 (tier 2.a), 0.76 (tier 2.b), 0.10 (tier 3); ChatGPT: 0.92, 0.49, 0.74, 0.05; Llama-2: 0.67, 0.16, -0.03, 0.02
Caveat
Tier 4 has no human annotations, so the trend cannot be extended to the most complex tier. The correlation metric captures judgment alignment but not actual leakage behaviour.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, ChatGPT, InstructGPT, Mixtral, Llama 2 / Llama 2 base Llama-2-Chat
Concepts
Failure mode
Datasets
Martin & Nissenbaum (2016) Vignettes [source]
Methods
Pearson correlation [eval]
Related work
FANTOM [context], Sap et al. 2022 (Neural Theory-of-Mind) [context]
Related findings
IC-1268, IC-1269, IC-1270
Extraction
automatic-extraction