Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
IOI dataset / IOI task
anchor
Note
no anchor recorded: the extractor captured no citation for this entry, and the citing paper's reference entry was never read. The paper does cite it, so this is a gap in our extraction rather than in the source
Findings
IC-015
Truncating MLP weights in Pythia-1b increases the probability of the correct answer in an Indirect Object Identification task
[eval]
IC-052
GPT-2 small's IOI circuit activations are linearly decomposable into features for the io, s, and pos attributes, with the l10h0 name mover's attention decomposing into sparse pairwise feature interactions
[eval]
IC-053
In GPT-2 small's l10h0 name mover queries, the io attribute is encoded with higher-magnitude features than the s attribute, and both are causally relevant, but SAEs preferentially learn io features due to the magnitude asymmetry
[eval]
IC-1250
GPT-2-medium shares 78% of its top attention heads between the IOI circuit and the colored objects circuit
[eval]
IC-1252
Circuit overlap between IOI and colored objects in GPT-2 decreases as model scale increases from medium to xl
[eval]
IC-167
Adding a PCA-derived control vector to the middle-layer residual stream improves logit-based reasoning accuracy on Pythia-1.4b, Pythia-2.8b, and Mistral-7B-Instruct
[eval]
IC-211
BLOOM-560M employs the same attention head circuit for indirect object identification in both English and Chinese
[eval]
IC-212
GPT-2-small and CPM-distilled converge on nearly identical IOI circuits despite being trained independently on English and Chinese
[eval]
IC-795
1D subspaces of MLP activations found by DAS in GPT-2 Small (IOI) and GPT-2 XL (factual recall) produce apparent causal effects that are interpretability illusions driven by causally disconnected components activating dormant pathways
[eval]
IC-809
Sparse dictionary features in Pythia-410m enable more precise causal localisation of indirect object identification behaviour than PCA, requiring fewer patches and smaller edit magnitudes for the same KL divergence
[eval]