IC-071DINOv2's dense features are dominated by global context, impairing fine-grained spatial detail

Congpei Qiu, Yanhao Wu, Wei Ke, Xiuxiu Bai, Tong Zhang

SourceRefining CLIP's Spatial Awareness: A Visual-Centric Perspective

The paper reports that DINOv2 produces dense feature artifacts that impair its ability to capture fine-grained details, resulting in abnormal representations dominated by global context. This is measured via unsupervised segmentation with CAUSE: DINOv2 achieves 29.9 mIoU on Cityscapes and 43.0 mIoU on COCO-Stuff. Affinity map visualizations (Figure 8) show that a query token's similarity map is spread broadly rather than focused on the local region, confirming the global-context dominance. The paper attributes this observation to Darcet et al. 2023.

Evidence
observational
Key metric
DINOv2 mIoU: 29.9 on Cityscapes, 43.0 on COCO-Stuff (Table 5)
Caveat
The interpretation of the artifacts as 'register tokens' is attributed to Darcet et al. 2023; this paper measures the performance but does not independently diagnose the mechanism.
Model
DINOv2
Concepts
Failure mode
Datasets
Cityscapes [eval], COCO-Stuff [eval]
Methods
CAUSE [eval]
Related work
Vision Transformers Need Registers [context]
Related findings
IC-069, IC-070
Extraction
automatic-extraction