The paper linearly probed or text-prompted 327 released DNNs on the 3D-PC benchmark. On the depth order task, 15 DNNs fell within the human accuracy confidence interval and 3 outperformed humans (74.73%). On VPT-basic, the best DNN (BEiT trained on ImageNet-21k) reached only 53.82% accuracy, while humans scored 86.82%. Commercial VLMs were at chance: ChatGPT4 52%, Gemini 52%, Claude 3 50%. The pattern is inverted relative to humans, for whom VPT-basic is easier than depth order.
Evidence
correlational
Key metric
Human depth order 74.73%, human VPT-basic 86.82%; best DNN VPT-basic (BEiT ImageNet-21k) 53.82%; ChatGPT4 52%, Gemini 52%, Claude 3 50% on VPT-basic; 15 DNNs within human CI on depth order, 3 outperformed humans
Caveat
The paper notes it could not isolate the precise monocular depth cues used by humans versus DNNs on the depth order task, so the strategies may differ even where accuracy matches.