SourceSparse autoencoders reveal selective remapping of visual concepts during adaptation
The authors train separate PatchSAEs on CLIP ViT-B/16's residual stream at layers 2, 5, 8, and 11. Qualitative inspection of reference images and segmentation masks shows that layer 11 latents represent semantic concepts (e.g., 'golden gate bridge'), layer 8 and 5 latents capture geometric shapes (e.g., 'triangle shape of the object'), and layer 2 latents are less interpretable, activated by local token attention rather than global semantics. The SAE training metrics also differ: L0 (active latents) ranges from 148.56 (layer 11) to 298.97 (layer 5).