IC-639The causal tracing pattern of MLP at early layers and attention at late layers is not stable across factual and syntactic phenomena in GPT-2 XL

Jingcheng Niu, Andrew Liu, Zining Zhu, Gerald Penn

SourceWhat does the Knowledge Neuron Thesis Have to do with Knowledge?

Meng et al. (2022) proposed that factual recall occurs in mid-layer MLPs and is copied to output by top-layer attention. The paper reproduces this causal tracing on syntactic phenomena (determiner-noun agreement, subject-verb agreement) and finds the clean two-site pattern does not hold. For subject-verb agreement, MLP modules show strongest causality at layers 30–35, later than attention (~25). Many factual traces also show MLP causality at the late site, contradicting the proposed division of labor.

Evidence
interventional
Key metric
For subject-verb agreement in GPT-2 XL, MLP has strongest causality at layers 30–35 vs attention at approximately layer 25
Caveat
The analysis is based on GPT-2 XL only. The authors note that the pattern is 'less stable' rather than entirely absent, and that the average indirect effect still shows some MLP/attention separation.
Model
GPT-2
Concepts
Depth-dependent structure
Datasets
BLIMP [eval], ParaRel [eval]
Methods
Causal Tracing [primary]
Related findings
IC-636, IC-637, IC-638
Extraction
automatic-extraction