Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
ImageNet-Segmentation
Findings
IC-447
CLIP ViT-B/16 produces noisy saliency maps and contains only 42 concept detectors, indicating poor visual interpretability
[eval]
IC-450
LLaVA with the original CLIP encoder produces noisy, non-sparse attention maps that poorly localize to the objects described in generated text
[eval]
IC-680
CLIP ViT's image token contributions are spatially localized to match described content, enabling zero-shot segmentation that outperforms existing CLIP-based methods
[eval]