IC-1472DINOv2 (86M parameters) achieves the smallest GEODE geographic disparity (2.46% Europe-Africa gap) among all 98 models in the testbed

Megan Richards, Polina Kirichenko, Diane Bouchacourt, Mark Ibrahim

SourceDoes Progress On Object Recognition Benchmarks Improve Generalization on Crowdsourced, Global Data?

Among 98 vision models spanning 16 architectures and 8 pretraining datasets, DINOv2, a self-supervised foundation model trained on auto-curated video data (LVD-142M), shows the smallest region performance disparity on GEODE at 2.46% between Europe and Africa subsets. While it still has a significant disparity on DollarStreet, its GEODE performance is remarkable for its 86M parameter size. The authors interpret this as evidence that data curation (auto-curated video data) offers a promising path to mitigating geographic disparities.

Evidence
correlational
Key metric
DINOv2 GEODE disparity: 2.46% (Europe-Africa accuracy difference); 86M parameters; trained on LVD-142M
Caveat
The authors note DINOv2 still has a significant region disparity on DollarStreet, and that the GEODE improvement may partly reflect the coarser class mapping (1-to-many) making GEODE easier for some models.
Model
DINOv2
Datasets
GEODE [eval], DollarStreet [eval]
Related work
Ramaswamy et al. 2023 (Beyond web-scraping: GEODE) [context]
Related findings
IC-1469, IC-1470, IC-1471
Extraction
automatic-extraction