IC-217AnyLoc's VLAD aggregation, learned unsupervised on the gallery, fails to generalise to out-of-distribution queries with large time gaps or seasonal changes

Issar Tzachor, Boaz Lerner, Matan Levy, Michael Green, Tal Berkovitz Shalev, Gavriel Habib, Dvir Samuel, Noam Korngut Zailer, Or Shimshi, Nir Darshan, Rami Ben-Ari

SourceEffoVPR: Effective Foundation Model Utilization for Visual Place Recognition

The paper reports that AnyLoc, which uses VLAD pooling learned in an unsupervised manner on the gallery, shows a significant performance drop on challenging VPR benchmarks. On Tokyo24/7 (day/night/sunset queries) AnyLoc achieves only 60.6 R@1, and on Nordland (seasonal changes) it drops to 16.1 R@1, well below DINOv2's zero-shot [cls] token (62.2 and 33.0 respectively). The authors attribute this to the VLAD dictionary overfitting to the gallery distribution and failing on out-of-distribution queries.

Evidence
correlational
Key metric
AnyLoc R@1: Pitts30k 87.7, Tokyo24/7 60.6, MSLs-val 68.7, Nordland 16.1
Caveat
The paper notes it implemented an online clustering scheme to handle AnyLoc's memory requirements, which may affect the exact numbers compared to the original publication.
Model
AnyLoc
Concepts
Failure mode
Datasets
Pitts30k [eval], Tokyo24/7 [eval], Nordland [eval]
Related findings
IC-215, IC-216
Extraction
automatic-extraction