The paper evaluates 9 released LLMs on the PersonalReddit dataset (520 profiles, 1066 labels across 8 attributes). GPT-4 achieves 85.5% top-1 and 95.2% top-3 accuracy, approaching human labeler performance. A clear scale trend is observed: Llama-2 7B reaches 51% while Llama-2 70B reaches 66%. The models achieve this at 100x lower cost and 240x lower time than human labelers, making large-scale privacy inference feasible.
Evidence
correlational
Key metric
GPT-4: 85.5% top-1, 95.2% top-3; Llama-2 7B: 51%; Llama-2 70B: 66%; 100x cost reduction, 240x time reduction vs. human labelers
Caveat
Human labelers had access to subreddit names and search engines, which models did not; the dataset is not publicly released due to privacy concerns