IC-257327 DNNs approach or exceed human accuracy on object depth order but are near chance on VPT-basic, while humans show the opposite pattern

Drew Linsley, Peisen Zhou, Alekh Karkada Ashok, Akash Nagaraj, Gaurav Gaonkar, Francis E Lewis, Zygmunt Pizlo, Thomas Serre

SourceThe 3D-PC: a benchmark for visual perspective taking in humans and machines

The paper linearly probed or text-prompted 327 released DNNs on the 3D-PC benchmark. On the depth order task, 15 DNNs fell within the human accuracy confidence interval and 3 outperformed humans (74.73%). On VPT-basic, the best DNN (BEiT trained on ImageNet-21k) reached only 53.82% accuracy, while humans scored 86.82%. Commercial VLMs were at chance: ChatGPT4 52%, Gemini 52%, Claude 3 50%. The pattern is inverted relative to humans, for whom VPT-basic is easier than depth order.

Evidence
correlational
Key metric
Human depth order 74.73%, human VPT-basic 86.82%; best DNN VPT-basic (BEiT ImageNet-21k) 53.82%; ChatGPT4 52%, Gemini 52%, Claude 3 50% on VPT-basic; 15 DNNs within human CI on depth order, 3 outperformed humans
Caveat
The paper notes it could not isolate the precise monocular depth cues used by humans versus DNNs on the depth order task, so the strategies may differ even where accuracy matches.
Model
BEiT, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Gemini, Claude 3, Stable Diffusion 2.0, MAE, DINOv2, SAM, MiDaS, Depth Anything
Concepts
Failure mode
Datasets
CO3D [source]
Methods
Linear Probing / Ridge regression linear probing / Linear probe / Linear probe fine-tuning / Linear regression probing / Linear ridge regression probes / Supervised probing / ERM linear probe [primary]
Related findings
IC-258, IC-259
Extraction
automatic-extraction