Light Dark How closely an explanation reproduces the behaviour of the model it explains, measured against the model's own outputs. It is a property of the explanation method and says nothing about whether the model itself is correct.
Findings FX-001 Pairwise Banzhaf interactions explain CLIP similarity more faithfully than single-score methods FX-002 FIXLIP gives SigLIP-2 higher pointing-game recognition than CLIP at ViT-B/32 and ViT-B/16 IC-002 Even when VLMs select the correct hypothesis, their explanations are often invalid or unhelpful IC-082 GPT-3.5, GPT-4o, Claude-3.5-Sonnet, and Llama-3.1-8B produce explanations on the BBQ social bias task that are systematically unfaithful for identity and behavior concepts while remaining faithful for context concepts, with specific patterns of hiding safety-measure influence and social bias IC-083 GPT-3.5, GPT-4o, and Claude-3.5-Sonnet produce unfaithful explanations on MedQA medical questions, omitting high-effect clinical concepts such as the patient's mental status while over-referencing low-effect concepts IC-1045 Factual GNN explanations do not capture the full data signal: retraining on explanations fails to reproduce predictions while retraining on residuals preserves them IC-1070 In Stable Diffusion v2.1, attention maps do not reliably predict the effect of prompt interventions on generated images, while conditional mutual information does IC-1071 Stable Diffusion v2.1's pixel-wise conditional mutual information localizes abstract words (adjectives, adverbs, verbs) more effectively than attention, but is less effective than attention for object segmentation IC-1341 DRUM's standard datalog rule extraction is fundamentally incomplete because its predictions depend on counting distinct rule-body matches IC-1408 OpenFlamingo and Idefics models generate low-quality explanations in zero-shot, but ICL and model scale significantly improve explanation CIDEr IC-1602 ResNet18, ResNet34, and MobileNetV2 pre-trained on CIFAR10 have decision functions well-approximated by a kernel machine using the trace NTK, with Kendall-τ correlations of 0.776, 0.786, and 0.700 IC-162 The majority of subject-verb agreement performance in Pythia-70m is explained by approximately 100 SAE feature nodes and in Gemma-2-2b by approximately 500 nodes, compared to approximately 1500 and 50000 neurons respectively IC-290 Zeroing out or doubling specific FFN neurons identified by the neuron path method causes significant accuracy changes in ViT and MAE models IC-373 In Pythia-2.8B, the specific attention heads and MLPs implementing retrieval depend on superficial input features, and request-patching preserves the natural retrieval mechanism IC-381 Individual knowledge is not parameter-localizable in GPT-J: existing localization methods (KN, ROME, KC) are neither faithful nor reliable IC-449 CLIP ViT-B/16 Grad-CAM explanations are highly sensitive to input noise, with SSIM dropping from 91.18% to 70.58% as noise standard deviation increases from 1/255 to 9/255 IC-458 Most VLM decoders show negative CC-SHAP on VALSE multiple-choice, indicating their explanations are less self-consistent than their answers, driven by a shift from text-dominant to image-dominant processing