IC-1227SAM's segmentation pretraining shifts attention heads toward local focus in deeper layers, unlike its MAE initialization which retains global attention throughout

Zihan Zhong, Zhiqiang Tang, Tong He, Haoyang Fang, Chun Yuan

SourceConvolution Meets LoRA: Parameter Efficient Finetuning for Segment Anything Model

The authors measure the mean attention distance (average distance between query patch position and attended locations, weighted by attention weights) for each of the 16 attention heads at each of the 25 layer depths, averaged over 500 randomly sampled images. In SAM, numerous heads in the deep blocks exhibit short mean attention distances, indicating a strong local attention pattern in later stages. In contrast, the MAE-pretrained ViT (SAM's initialization) displays consistently long mean attention distances across all layers. This demonstrates that SAM's large-scale segmentation pretraining transforms the ViT's attention from global-oriented to local-oriented, with the effect concentrated in deeper layers.

Evidence
correlational
Model
SAM, MAE
Concepts
Depth-dependent structure
Related findings
IC-1226
Extraction
automatic-extraction