When the authors built attribution heatmaps with Integrated Gradients, nearly every map lit up the same top and bottom left corner patches no matter what the image showed. Those corners are most likely the model reusing spare tokens for internal bookkeeping rather than looking at anything meaningful, and they dominate the clustering step so heavily that you cannot draw conclusions from these maps without switching to a better attribution method first.
Evidence
observational
Caveat
The authors give the register-token explanation as a likely interpretation, not as a tested claim.