SourceLabel-Focused Inductive Bias over Latent Object Features in Visual Classification
The paper extracts class-level feature centroids from pre-trained ViT, MAE, and ResNet50 (both supervised and contrastive/MoCo v2) on ImageNet-1K. In all four models, classes that share visual appearance but differ semantically—mop and broom (shared stick), puck and crutch (shared cylindrical shape)—have centroids that are closely located in the final-layer feature space. This indicates the representations are dominated by input-domain visual similarity rather than by the semantic structure of the labels. The effect persists under contrastive pre-training (MoCo v2), which only slightly decouples some affected pairs but still fails to separate broom from puck.