The authors define successor heads as attention heads whose effective OV circuit assigns higher logit to the successor token than to any other token in the same ordinal sequence for more than half of the succession dataset tokens. They measure the best successor score (proportion of tokens incremented) across models of three architectures and nearly three orders of magnitude in parameter count. Successor heads are found in all tested models, with the best score generally increasing with model size. The phenomenon is observed across eight token tasks: numbers, number words, cardinal words, days, day prefixes, months, month prefixes, and letters.
Evidence
correlational
Key metric
successor heads form in llms with as few as 31 million parameters, and at least as many as 12 billion parameters, such as gpt-2, pythia, and llama-2
Caveat
The original llama-7b does not generalize well to the held-out roman numeral task, suggesting the successor head mechanism may be less robust in that model for out-of-distribution ordinal sequences.