Light Dark A result that comes from the apparatus used to study the model rather than from the model: a pattern dominating an attribution map because of how the map is computed, or a comparison coming out one way because the metric shares an assumption with one of the methods it ranks. Attribution and evaluation are the two places this shows up and they are one concept, because the error being named is the same in both -- reading a property of the instrument as a property of the thing measured.
Findings IC-053 In GPT-2 small's l10h0 name mover queries, the io attribute is encoded with higher-magnitude features than the s attribute, and both are causally relevant, but SAEs preferentially learn io features due to the magnitude asymmetry IC-054 Gemma 2 2B performance degrades substantially when routed through Gemma Scope SAEs, and SAE-based feature suppression causes broad cross-domain degradation rather than targeted knowledge removal IC-075 GPT-4o-mini and Qwen2.5-72B achieve low precision and recall when used as API retrievers for workflow orchestration IC-1069 Stable Diffusion v2.1's compositional understanding on the ARO benchmark is significantly higher than previously reported by MMSE-based scoring IC-1380 Pythia and OPT small models exhibit non-trivial performance on BigBench tasks that is invisible under beam search but revealed by extensive random sampling IC-147 CLIP's OOD performance on rendition domains is largely an artifact of domain contamination in its web-scale training data IC-208 Llama-3-8B-Inst-SimPO does not outperform Llama-3-70B-Inst on WildBench, contrary to its advantage on AlpacaEval-2.0, but performs comparably on information-seeking and creative tasks IC-209 LLM judges (GPT-3.5-turbo-1106, GPT-4o-mini, GPT-4o, Claude-3-5-sonnet) implicitly prioritize style over factuality and safety when scoring pairwise preferences IC-295 The ViT model's ECE can be reduced to near-zero by trivial mean-replacement recalibration while maintaining test accuracy, but NLL increases from 65.35 to 144.66, demonstrating that ECE and accuracy alone are an insufficient reporting standard for calibration IC-297 Applying tuned-lens transformations to the residual stream decreases the apparent multi-layer SAE feature activity from 54–88% to 37–41% of total variance IC-385 TAR-bio-v1 retains bio-weaponization knowledge despite appearing to unlearn it; a different prompt template and answer extraction method reveals accuracy above 45% on WMDP-bio IC-409 Knowledge editing methods correct verified hallucinations in Llama2-7B, Llama3-8B, and Mistral-v0.3-7B far less effectively than their scores on existing benchmarks suggest IC-427 Newer base models (post-November 2023) outperform older ones by 7.3 points on MMLU and 19.1 points on GSM8K controlling for pretraining compute, but this gap vanishes after fine-tuning all models on the same task-relevant data IC-428 Qwen 1.5 appears to Pareto-dominate Pythia and LLaMA 2 on MMLU and GSM8K, but after adjusting for test task training all three model families exhibit equivalent scaling IC-467 Llama-3.1-405B's standard speculative decoding verification rejects correct continuations from GPT-4o, Llama-3.1-8B, and human text, accepting only roughly two tokens before the first rejection for GPT-4o IC-795 1D subspaces of MLP activations found by DAS in GPT-2 Small (IOI) and GPT-2 XL (factual recall) produce apparent causal effects that are interpretability illusions driven by causally disconnected components activating dormant pathways IC-854 The apparent self-correction improvement in constrained generation (Madaan et al., 2023) is an artefact of a sub-optimal initial prompt, not a genuine model capability TM-002 The mean-difference concept vector scores higher than trained classifiers under cosine similarity TM-010 Fixed corner patches dominate Integrated Gradients maps regardless of image content