Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Understanding In-Context Learning in Transformers and LLMs by Learning to Learn Discrete Functions
2024-01-16
· ICLR 2024 oral ·
anchor
Findings
IC-1219
Frozen GPT-2 XL contains pre-trained attention heads that implement the nearest-neighbor algorithm
IC-1220
GPT-4, GPT-3.5-turbo, and Llama-2-70B can implement learning algorithms in-context on novel boolean functions, competing with nearest-neighbor baselines
IC-1221
LLM performance on in-context boolean function learning is scale-dependent, with GPT-2 failing and Llama-2 models improving gradually with size