Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Interpreting CLIP's Image Representation via Text-Based Decomposition
2024-01-16
· ICLR 2024 oral ·
anchor
Findings
IC-677
CLIP ViT's image representation is primarily constructed by the last 4 MSA layers, with MLPs and early MSA layers contributing negligibly
IC-678
Specific attention heads in CLIP ViT-L's last 4 layers encode specific image properties (color, shape, location, counting, texture) that are linearly recoverable via text directions
IC-679
CLIP relies on background/location as a spurious cue for bird classification, and ablating geolocation heads improves worst-group accuracy by 25.2%
IC-680
CLIP ViT's image token contributions are spatially localized to match described content, enabling zero-shot segmentation that outperforms existing CLIP-based methods