IC-496GPT-4o achieves the highest location inference accuracy among evaluated models, reaching 98.16% for country, 60.23% for city, and 27.13% for zip code from street view images

Chejian Xu, Jiawei Zhang, Zhaorun Chen, Chulin Xie, Mintong Kang, Yujin Potter, Zhun Wang, Zhuowen Yuan, Alexander Xiong, Zidi Xiong, Chenhui Zhang, Lingzhi Yuan, Yi Zeng, Peiyang Xu, Chengquan Guo, Andy Zhou, Jeffrey Ziwei Tan, Xuandong Zhao, Francesco Pinto, Zhen Xiang, Yu Gai, Zinan Lin, Dan Hendrycks, Bo Li, Dawn Song

SourceMMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models

The paper evaluates image-to-text models on a dataset of 1,816 Google Street View images collected after 2023, asking models to infer location at four granularities (country, state, city, zip code). GPT-4o outperforms all other models at every granularity. Notably, existing models rarely refuse to predict locations, indicating a lack of awareness of location privacy risks. GPT-4V is the only model that occasionally refuses (within 1.61%).

Evidence
correlational
Key metric
gpt-4o: 98.16% country, 75.40% state, 60.23% city, 36.55% zip code range, 27.13% zip code; gpt-4v: 91.03% country, 44.60% state, 40.00% city, 17.47% zip code range, 12.18% zip code
Caveat
The dataset includes only images taken after 2023 to avoid memorization, but models with large knowledge bases may still have seen some of these locations through other channels.
Model
GPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4V / GPT-4 vision, Llama-3-2-Vision Llama-3.2-90B-Vision-Instruct, Nova Lite, Gemini 1.5 / Gemini Pro 1.5
Concepts
Failure mode
Related findings
IC-495, IC-497, IC-498
Extraction
automatic-extraction