IC-217AnyLoc's VLAD aggregation, learned unsupervised on the gallery, fails to generalise to out-of-distribution queries with large time gaps or seasonal changes
Issar Tzachor, Boaz Lerner, Matan Levy, Michael Green, Tal Berkovitz Shalev, Gavriel Habib, Dvir Samuel, Noam Korngut Zailer, Or Shimshi, Nir Darshan, Rami Ben-Ari
The paper reports that AnyLoc, which uses VLAD pooling learned in an unsupervised manner on the gallery, shows a significant performance drop on challenging VPR benchmarks. On Tokyo24/7 (day/night/sunset queries) AnyLoc achieves only 60.6 R@1, and on Nordland (seasonal changes) it drops to 16.1 R@1, well below DINOv2's zero-shot [cls] token (62.2 and 33.0 respectively). The authors attribute this to the VLAD dictionary overfitting to the gallery distribution and failing on out-of-distribution queries.
The paper notes it implemented an online clustering scheme to handle AnyLoc's memory requirements, which may affect the exact numbers compared to the original publication.