After anonymizing all PersonalReddit comments with Azure Language Service (removing locations, persons, organizations, dates, ages, numbers, currencies), GPT-4's accuracy decreases but remains substantial. Location accuracy drops from approximately 86% to approximately 55%. The anonymization effect diminishes with hardness: 41.1% accuracy decrease at hardness 1 but only 7.7% at hardness 5, indicating the model relies on subtle contextual and linguistic cues that anonymizers do not remove.
Evidence
correlational
Key metric
GPT-4 location: ~86% to ~55% after anonymization; accuracy decrease by hardness: 41.1% (h1), 19.0% (h2), 13.6% (h3), 31.6% (h4), 7.7% (h5)
Caveat
Only 5 of 8 attributes were supported by the anonymizer (location, age, occupation, place of birth, income); the anonymizer threshold was set to 0.4