Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
2024-01-16
· ICLR 2024 spotlight ·
anchor
Findings
IC-1015
GPT-J and 10 other LLMs exhibit overthinking: calibrated accuracy given incorrect few-shot demonstrations peaks at a critical layer then declines, and ablating 5 false induction heads in late layers reduces the accuracy gap by 38.9% on average