Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Boosting the visual interpretability of CLIP via adversarial fine-tuning
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-447
CLIP ViT-B/16 produces noisy saliency maps and contains only 42 concept detectors, indicating poor visual interpretability
IC-448
CLIP ViT-L/14 achieves 0% accuracy under 2/255 and 4/255 L-infinity adversarial perturbations across all 15 evaluation datasets
IC-449
CLIP ViT-B/16 Grad-CAM explanations are highly sensitive to input noise, with SSIM dropping from 91.18% to 70.58% as noise standard deviation increases from 1/255 to 9/255
IC-450
LLaVA with the original CLIP encoder produces noisy, non-sparse attention maps that poorly localize to the objects described in generated text