Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Unveiling and Manipulating Prompt Influence in Large Language Models
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1127
The LM head in GPT-2, GPT-J, BLOOM, Pythia, and LLaMA-2 projects all input token hidden states into interpretable token distributions over the vocabulary, and these distributions converge approximately monotonically toward the final layer's distribution