IC-1470Geographic disparities (Europe-Africa accuracy gap) are large across all 98 models and have more than tripled between least and best performing models on DollarStreet

Megan Richards, Polina Kirichenko, Diane Bouchacourt, Mark Ibrahim

SourceDoes Progress On Object Recognition Benchmarks Improve Generalization on Crowdsourced, Global Data?

The paper measures the maximum absolute accuracy difference between any two regions (geographic disparity) for each of 98 models on DollarStreet and GEODE. All models show substantial disparities: ResNet models average 14.5% on DollarStreet and 5.0% on GEODE, while the best-performing CLIP model shows 17.0% on DollarStreet and 6.5% on GEODE. As ImageNet accuracy increases across the model testbed, the Europe-Africa gap on DollarStreet widens, with disparities more than tripling between the least and most performant models. Europe improves at almost double the rate of Africa.

Evidence
correlational
Key metric
ResNet average disparity: 14.5% (DollarStreet), 5.0% (GEODE); best CLIP disparity: 17.0% (DollarStreet), 6.5% (GEODE); disparities more than tripled between least and best models on DollarStreet; Europe improves at almost double the rate of Africa
Caveat
The disparity is measured as the max absolute difference between any two regions, which is sensitive to the specific regions present in the dataset. The paper notes GEODE shows much more similar rates of improvement across regions compared to DollarStreet.
Model
CLIP / CLIP-ViT (LC), ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN, DINOv2, ViT, ConvNeXt, RegNet, MobileNetV3, VGG / VGG13, HRNet, FLAVA, EVA-CLIP, MLP-Mixer, EdgeNeXt, ReXNet
Datasets
DollarStreet [eval], GEODE [eval], ImageNet-1k / ImageNet / ImageNet-1k-val / ImageNet-Val [eval]
Related findings
IC-1469, IC-1471, IC-1472
Extraction
automatic-extraction