Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Grad-CAM
Findings
IC-044
Tulu-2-13B's internal activations contain a linearly decodable, faithful representation of input-context propositions that persists under prompt injection and backdoor attacks where outputs become unfaithful
[supporting]
IC-1156
The VPT advantage over FT for ViT-B/16 is not explained by overfitting resistance or additional optimization dimensions; the specific feature-preservation mechanism of VPT is the key factor
[supporting]
IC-447
CLIP ViT-B/16 produces noisy saliency maps and contains only 42 concept detectors, indicating poor visual interpretability
[primary]
IC-449
CLIP ViT-B/16 Grad-CAM explanations are highly sensitive to input noise, with SSIM dropping from 91.18% to 70.58% as noise standard deviation increases from 1/255 to 9/255
[primary]
IC-644
CLIP relies on spurious features (background) as a shortcut in zero-shot classification, and conditioning on the correct background reduces this reliance
[primary]
IC-680
CLIP ViT's image token contributions are spatially localized to match described content, enabling zero-shot segmentation that outperforms existing CLIP-based methods
[compared-to]
IC-742
ResNet-50-BN on Waterbirds relies on background as a spurious feature for classification, and this shortcut is invisible to entropy-based confidence metrics
[primary]