Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
LIMA
anchor
Findings
IC-1576
Base and aligned LLMs share 77.7% of top-1 token predictions, with distribution shifts concentrated in stylistic tokens rather than knowledge content
[context]
IC-1577
Base LLMs prompted with URiAL (3 restyled in-context examples + system prompt) match or surpass their SFT/RLHF-aligned counterparts on multi-aspect evaluation
[context]
IC-1577
Base LLMs prompted with URiAL (3 restyled in-context examples + system prompt) match or surpass their SFT/RLHF-aligned counterparts on multi-aspect evaluation
[source]
IC-986
Most LLMs lack tool usage awareness, with only ChatGPT exceeding 70% F1 in zero-shot evaluation
[eval]