Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Interpreting the Second-Order Effects of Neurons in CLIP
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-150
Second-order effects of CLIP's MLP neurons are concentrated in late layers (8–10 of 12 in ViT-B/32)
IC-151
Each CLIP neuron's second-order effect is approximately a single linear direction in the joint text-image space, significant for fewer than 2% of images
IC-152
CLIP's polysemantic neurons encode spurious correlations between unrelated concepts that can be exploited to generate adversarial misclassifications