Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1576
Base and aligned LLMs share 77.7% of top-1 token predictions, with distribution shifts concentrated in stylistic tokens rather than knowledge content
IC-1577
Base LLMs prompted with URiAL (3 restyled in-context examples + system prompt) match or surpass their SFT/RLHF-aligned counterparts on multi-aspect evaluation