Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-866
A prefix applied to Llama-7B's first attention layer preserves the relative attention distribution over content positions and only adds a constant-direction bias to the attention block output
IC-867
In GPT-2 prefix-tuned on the emotion dataset, attention over prefix positions is nearly constant across inputs, collapsing the effective bias subspace to a single direction in most layers