The paper compares k-means clustering on DINOv2's middle-layer features with CLIP's under identical settings. DINOv2's clustering results appear smoother with larger contiguous blocks sharing similar semantics, but they do not delineate individual object boundaries. On COCO val2017, DINOv2 achieves only AP 0.8 (class-agnostic) and 0.2 (class-aware) for instance segmentation, far below SAM and CLIP-based methods. The authors attribute this to DINOv2's self-supervised pretraining objective, which encourages global semantic coherence rather than instance-level discrimination.