Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Successor Heads: Recurring, Interpretable Attention Heads In The Wild
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1369
Successor heads that increment ordinal-sequence tokens exist in Pythia, GPT-2, and Llama-2 models from 31M to 12B parameters
IC-1370
MLP0 representations of ordinal-sequence tokens in Pythia-1.4b contain linearly decodable mod-10 features that are causally important for incrementation
IC-1371
The Pythia-1.4b successor head l12h0 exhibits interpretable polysemanticity, performing successorship, acronym prediction, copying, and greater-than behaviors on natural language data
IC-1372
Successor heads in Pythia-1.4b exhibit a greater-than bias: the OV circuit assigns systematically higher logits to tokens with greater ordinal values than the input, impairing decrementation