Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Synergy and Diversity in CLIP: Enhancing Performance Through Adaptive Backbone Ensembling
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-318
CLIP backbones from different architectures (ViTs and ResNets) trained with the same data and objective exhibit complementary strengths, with an oracle per-image backbone selection improving zero-shot accuracy by up to 43.5% over the best single backbone
IC-319
Different CLIP backbones exhibit distinct robustness profiles to specific image perturbations, with each architecture being most resilient to a different transformation