Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-215
DINOv2's zero-shot attention maps focus on irrelevant foreground objects (vehicles, advertisements) rather than scene structure, degrading its VPR recall on challenging datasets
IC-216
DINOv2's value (V) facet from self-attention at layer n-1 encodes the most effective local features for VPR re-ranking, outperforming query and key facets, and layer n-1 outperforms the final layer n
IC-217
AnyLoc's VLAD aggregation, learned unsupervised on the gallery, fails to generalise to out-of-distribution queries with large time gaps or seasonal changes