Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
ImageNet-R / ImageNet-Rendition
anchor
Findings
IC-1469
Progress on standard ImageNet generalization benchmarks is 2.5x faster than progress on crowdsourced global data (DollarStreet, GEODE) across 98 vision models
[eval]
IC-147
CLIP's OOD performance on rendition domains is largely an artifact of domain contamination in its web-scale training data
[eval]
IC-1496
A single frozen transformer block from LLaMA-7B consistently improves performance across diverse visual tasks when appended to existing visual encoders
[eval]
IC-150
Second-order effects of CLIP's MLP neurons are concentrated in late layers (8–10 of 12 in ViT-B/32)
[eval]
IC-151
Each CLIP neuron's second-order effect is approximately a single linear direction in the joint text-image space, significant for fewer than 2% of images
[eval]
IC-1520
OpenCLIP's per-sample zero-shot accuracy on ImageNet-based OOD benchmarks is strongly correlated with the perceptual similarity between that sample and its nearest neighbor in LAION-400M
[eval]
IC-448
CLIP ViT-L/14 achieves 0% accuracy under 2/255 and 4/255 L-infinity adversarial perturbations across all 15 evaluation datasets
[eval]
IC-643
Conditioning CLIP on correct contextual attributes in the text prompt improves zero-shot classification accuracy across 13 image transformations
[eval]