SourceDistributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers
The paper evaluates the effect of Layer-Selective Rank Reduction (LASER) on the performance of Pythia models on an Indirect Object Identification (IOI) task. They measure the average probability of predicting the correct indirect object token ('[io]'), the distractor subject token ('[s]'), and the generic token 'the' over a dataset of 100 IOI sentences. Applying LASER to the MLP weights of Pythia-1b boosts the probability ratio of '[io]' over 'the' from 2.3x to 12.3x at 14k training steps.