SourceConvolution Meets LoRA: Parameter Efficient Finetuning for Segment Anything Model
The authors measure the mean attention distance (average distance between query patch position and attended locations, weighted by attention weights) for each of the 16 attention heads at each of the 25 layer depths, averaged over 500 randomly sampled images. In SAM, numerous heads in the deep blocks exhibit short mean attention distances, indicating a strong local attention pattern in later stages. In contrast, the MAE-pretrained ViT (SAM's initialization) displays consistently long mean attention distances across all layers. This demonstrates that SAM's large-scale segmentation pretraining transforms the ViT's attention from global-oriented to local-oriented, with the effect concentrated in deeper layers.