Among 98 vision models spanning 16 architectures and 8 pretraining datasets, DINOv2, a self-supervised foundation model trained on auto-curated video data (LVD-142M), shows the smallest region performance disparity on GEODE at 2.46% between Europe and Africa subsets. While it still has a significant disparity on DollarStreet, its GEODE performance is remarkable for its 86M parameter size. The authors interpret this as evidence that data curation (auto-curated video data) offers a promising path to mitigating geographic disparities.
The authors note DINOv2 still has a significant region disparity on DollarStreet, and that the GEODE improvement may partly reflect the coarser class mapping (1-to-many) making GEODE easier for some models.