SourceQPM: Discrete Optimization for Globally Interpretable Image Classification
The paper uses CLIP ViT-L/14 text embeddings as a proxy for ground-truth class similarity on ImageNet-1k, computing cosine similarities between the text embeddings of all 1000 class names. Fine-grained visual similarities are poorly captured: the pair rottweiler and doberman, which are visually very similar dog breeds, ranks at position 828 out of all pairs. It is ranked behind semantically unrelated pairs such as hog and tank, lemon and yawl, or hamster and snail. More commonly used terms (e.g., orangutan and gorilla) are correctly associated, but fine-grained distinctions are not. The authors had to set the number of similar classes to consider to 1250 to include the rottweiler/doberman pair.