anchor
Findings
- IC-528The knowledge localization assumption fails for a large fraction of facts in GPT-2, Llama2-7B, and Llama3-8B, with 77% of facts classified as inconsistent knowledge in Llama3-8B [eval]
- IC-529For inconsistent knowledge in GPT-2, Llama2-7B, and Llama3-8B, the knowledge neurons are associated with the specific query rather than the fact, as shown by differential effects of suppressing or enhancing query-specific versus neighbor neurons [eval]
- IC-530The attention module in GPT-2, Llama2-7B, and Llama3-8B plays a selective role in knowledge expression by activating specific knowledge neurons for a given query, as demonstrated by suppressing or enhancing attention scores at knowledge synapse positions [eval]
- IC-636Syntactic phenomena (determiner-noun and subject-verb agreement) localize to the same topmost-layer MLP neurons as factual information in BERT, GPT-2, and Llama-2 [eval]
- IC-638ROME editing on GPT-2 XL and Llama-2 7B achieves high reliability but fails under bijective symmetry (23.71%–33.64%) and synonymous invariance (52.35%–58.36%) criteria [eval]
- IC-639The causal tracing pattern of MLP at early layers and attention at late layers is not stable across factual and syntactic phenomena in GPT-2 XL [eval]