IC-1469Progress on standard ImageNet generalization benchmarks is 2.5x faster than progress on crowdsourced global data (DollarStreet, GEODE) across 98 vision models
Megan Richards, Polina Kirichenko, Diane Bouchacourt, Mark Ibrahim
The paper measures the slope of accuracy improvement on each benchmark as a function of ImageNet accuracy across 98 released vision models spanning 16 architectures. Standard OOD benchmarks (ImageNet-V2, ImageNet-Sketch, ImageNet-Rendition, ObjectNet) improve by 62.75% on average, while DollarStreet improves by only 18.92% and GEODE by 33.5%. The linear fits are statistically significant with high R-squared values, confirming the progress gap is not noise. The gap is consistent across both geographic datasets despite their different curation and class sets.
Evidence
correlational
Key metric
OOD average net improvement +62.75% (progress rate 1.44) vs DollarStreet +18.92% (progress rate 0.53, R²=0.93) and GEODE +33.5%; progress gap 2.5x; individual benchmarks: ImageNet-V2 +37.74% (2.2x), ImageNet-Sketch +63.00% (2.6x), ImageNet-Rendition +73.42% (2.8x), ObjectNet +51.84% (2.8x)
Caveat
The progress gap relies on a linear fit of accuracy trends; the authors note the fits are well supported by high R² values but this is an assumption about the shape of the relationship.