In tier 4, models generate meeting summaries or personal action items from a transcript where a secret about person X was discussed before X joined. GPT-4 leaks the secret in 39% of summaries and 29% of action items; ChatGPT leaks in 57% and 38% respectively. All models also frequently omit public information (GPT-4: 10% in summaries, 76% in action items). The authors hypothesize the summary task is harder because the model must reason that X is among the recipients. Manual inspection found 16 additional nuanced leaks in GPT-4 action items not caught by string matching.
Only 20 transcripts were used. Scenarios were generated by GPT-4 (familiarity concern); a control with ChatGPT-generated scenarios confirmed GPT-4 still outperforms ChatGPT. String-match evaluation underestimates true leakage.