Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Beyond Memorization: Violating Privacy via Inference with Large Language Models
2024-01-16
· ICLR 2024 spotlight ·
anchor
Findings
IC-1361
GPT-4 and other state-of-the-art LLMs achieve near-human accuracy in inferring personal attributes from unstructured text
IC-1362
State-of-the-art text anonymization is insufficient to prevent GPT-4 from inferring personal attributes
IC-1363
Current model alignment does not filter privacy-invasive prompts across major LLM providers
IC-1364
GPT-4 can extract personal information from users through adversarial chatbot conversations