IC-1361GPT-4 and other state-of-the-art LLMs achieve near-human accuracy in inferring personal attributes from unstructured text

Robin Staab, Mark Vero, Mislav Balunovic, Martin Vechev

SourceBeyond Memorization: Violating Privacy via Inference with Large Language Models

The paper evaluates 9 released LLMs on the PersonalReddit dataset (520 profiles, 1066 labels across 8 attributes). GPT-4 achieves 85.5% top-1 and 95.2% top-3 accuracy, approaching human labeler performance. A clear scale trend is observed: Llama-2 7B reaches 51% while Llama-2 70B reaches 66%. The models achieve this at 100x lower cost and 240x lower time than human labelers, making large-scale privacy inference feasible.

Evidence
correlational
Key metric
GPT-4: 85.5% top-1, 95.2% top-3; Llama-2 7B: 51%; Llama-2 70B: 66%; 100x cost reduction, 240x time reduction vs. human labelers
Caveat
Human labelers had access to subreddit names and search engines, which models did not; the dataset is not publicly released due to privacy concerns
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Llama 2 / Llama 2 base, Claude Instant 1, PaLM 2
Datasets
PAN 2018 [eval], ACS Income [eval]
Related work
Hegselmann et al. (TabLLM) [context]
Related findings
IC-1362, IC-1363, IC-1364
Extraction
automatic-extraction