Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Refining CLIP's Spatial Awareness: A Visual-Centric Perspective
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-069
EVA-CLIP's dense patch features are semantically contaminated by surrounding context, degrading their spatial quality
IC-070
Region-language alignment fine-tuning degrades EVA-CLIP's spatial awareness as measured by unsupervised segmentation
IC-071
DINOv2's dense features are dominated by global context, impairing fine-grained spatial detail