Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
2024-01-16
· ICLR 2024 spotlight ·
anchor
Findings
IC-1267
LLMs' alignment with human privacy judgments drops sharply as contextual complexity increases from tier 1 to tier 3
IC-1268
LLMs leak private information in theory-of-mind scenarios even when explicitly instructed to preserve privacy
IC-1269
LLMs leak secrets to inappropriate recipients in meeting summarization and action-item generation tasks
IC-1270
Chain-of-thought prompting does not mitigate privacy leakage in GPT-4 or ChatGPT