Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Convolution Meets LoRA: Parameter Efficient Finetuning for Segment Anything Model
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1226
SAM's ViT-B encoder achieves 54.2% ImageNet-1k linear probing accuracy versus 67.7% for MAE's ViT-B, indicating its segmentation pretraining impairs high-level semantic representation
IC-1227
SAM's segmentation pretraining shifts attention heads toward local focus in deeper layers, unlike its MAE initialization which retains global attention throughout