IC-258DNN accuracy on 3D perception tasks correlates with ImageNet object classification accuracy, suggesting 3D cues emerge as a byproduct of object recognition training

Drew Linsley, Peisen Zhou, Alekh Karkada Ashok, Akash Nagaraj, Gaurav Gaonkar, Francis E Lewis, Zygmunt Pizlo, Thomas Serre

SourceThe 3D-PC: a benchmark for visual perspective taking in humans and machines

Across the 317 TIMM models, the paper measured the correlation between each model's ImageNet classification accuracy and its performance on the 3D-PC tasks. Depth order accuracy correlated strongly with ImageNet accuracy (rho = 0.66, p < 0.001), and VPT-basic accuracy showed a weaker but significant correlation (rho = 0.34, p < 0.001). The difference between the two correlations was itself significant (rho = 0.32, p < 0.001), indicating that the 3D cues that emerge with scale are well-suited for depth ordering but poorly suited for perspective taking.

Evidence
correlational
Key metric
rho = 0.66 (depth order vs ImageNet, p < 0.001); rho = 0.34 (VPT-basic vs ImageNet, p < 0.001); difference in correlations rho = 0.32, p < 0.001
Caveat
The authors note that more work is needed to identify a causal relationship between the development of monocular depth cues and object recognition accuracy; the correlation is not evidence of causation.
Model
BEiT, Swin Transformer
Concepts
Scale-dependent behaviour
Datasets
ImageNet-1k / ImageNet / ImageNet-1k-val / ImageNet-Val [eval]
Methods
Linear Probing / Ridge regression linear probing / Linear probe / Linear probe fine-tuning / Linear regression probing / Linear ridge regression probes / Supervised probing / ERM linear probe [primary]
Related findings
IC-257, IC-259
Extraction
automatic-extraction