The paper measures Pearson correlation between model and human privacy judgments across four tiers of increasing contextual complexity. In tier 1 (simple sensitivity rating), GPT-4 achieves 0.86 and ChatGPT 0.92 correlation. By tier 3 (theory-of-mind scenarios with three parties), GPT-4 drops to 0.10 and ChatGPT to 0.05. All six evaluated models show the same downward trend, with open-source models (Llama-2, Mixtral) starting lower and ending near zero. The authors attribute this to the models' inability to reason about information flow in complex social contexts.
Tier 4 has no human annotations, so the trend cannot be extended to the most complex tier. The correlation metric captures judgment alignment but not actual leakage behaviour.