Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Linearity of Relation Decoding in Transformer Language Models
2024-01-16
· ICLR 2024 spotlight ·
anchor
Findings
IC-1553
GPT-J, GPT-2-XL, and Llama-13B decode approximately 48% of tested relations via a linear transformation on the subject representation, and this structure causally influences predictions
IC-1554
LRE faithfulness in GPT-J is concentrated in intermediate layers and drops sharply in later layers, consistent with a mode switch from relational encoding to next-token prediction
IC-1555
GPT-J's internal representations contain correct factual knowledge even when the model outputs falsehoods under repetition or instruction distraction prompts