Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
CUB-200-2011 / CUB200
Findings
IC-1165
ProtoPFormer achieves 42.2% attribute identification accuracy on CUB-200-2011 in a 7-rater human evaluation
[eval]
IC-280
CLIP, OpenCLIP, and SigLIP exhibit intra-modal misalignment: intra-modal similarity comparisons are suboptimal for image-to-image and text-to-text retrieval
[eval]
IC-281
SLIP's intra-modal self-supervised loss reduces intra-modal misalignment, making inter-modal inversion unnecessary for image retrieval
[eval]
IC-643
Conditioning CLIP on correct contextual attributes in the text prompt improves zero-shot classification accuracy across 13 image transformations
[eval]
IC-737
CLIP ViT-B/16 binarized dot products yield 0.50–0.58 accuracy on binary concept presence queries across five image classification datasets
[eval]
IC-738
BLIP-2 ViT-G FlanT5XL achieves 0.70–0.87 zero-shot accuracy on binary concept presence queries, competitive on most datasets but weaker on fine-grained CUB-200
[eval]
IC-739
GPT-3.5-turbo-0613 combined with CLIP produces more faithful concept-salience pseudo-labels than LLaMA-2-13B-Chat, InstructBLIP, or LLaVA-1.5B on most of five datasets
[eval]
IC-984
CLIP ViT-B/16's representation space does not reliably preserve semantic similarity as measured by shared image tags
[eval]