Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Super NaturalInstructions
anchor
Findings
IC-991
LLaMA-2, Falcon-7B, and GPT-3.5-Turbo exhibit large performance spread (up to 76 accuracy points) across semantically equivalent prompt formats, and model comparison rankings are frequently reversed by format choice
[eval]
IC-992
LLaMA-2-7B's last hidden layer encodes the prompt format with high identifiability, and the separability of format embeddings in the top two principal components correlates with performance spread
[eval]