SourceEffoVPR: Effective Foundation Model Utilization for Visual Place Recognition
The paper visualises DINOv2's self-attention maps on VPR query images and shows that the pre-trained model's strongest attention is drawn to dynamic or irrelevant foreground elements such as vehicles, pedestrians, and temporary advertisements, rather than to the building structures and scene layout that determine location. This misdirected attention limits DINOv2's zero-shot VPR performance, particularly on datasets with large viewpoint or illumination changes. The effect is illustrated across multiple examples (Figures 3b, 4a, S3) and is quantified by the gap between DINOv2's zero-shot recall on easier (Pitts30k: 78.1) versus harder (Tokyo24/7: 62.2, Nordland: 33.0) benchmarks.