IC-420CLIP ViT-B/16's SAE latent interpretability is depth-dependent: layer 11 encodes semantic object concepts while layers 2, 5, and 8 encode local shapes and attention patterns

Hyesu Lim, Jinho Choi, Jaegul Choo, Steffen Schneider

SourceSparse autoencoders reveal selective remapping of visual concepts during adaptation

The authors train separate PatchSAEs on CLIP ViT-B/16's residual stream at layers 2, 5, 8, and 11. Qualitative inspection of reference images and segmentation masks shows that layer 11 latents represent semantic concepts (e.g., 'golden gate bridge'), layer 8 and 5 latents capture geometric shapes (e.g., 'triangle shape of the object'), and layer 2 latents are less interpretable, activated by local token attention rather than global semantics. The SAE training metrics also differ: L0 (active latents) ranges from 148.56 (layer 11) to 298.97 (layer 5).

Evidence
observational
Caveat
The depth-dependent interpretability finding is qualitative (based on reference image inspection) and reported in the appendix; no quantitative interpretability metric is provided to compare layers.
Model
CLIP / CLIP-ViT (LC)
Concepts
Depth-dependent structure
Datasets
ImageNet-1k / ImageNet / ImageNet-1k-val / ImageNet-Val [train]
Related findings
IC-419
Extraction
automatic-extraction