Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-099
GPT-4 and Claude 3 Opus can be prompted to selectively underperform on WMDP while maintaining general performance on MMLU and CSQA
IC-100
GPT-4, GPT-3.5, Claude 3, Llama 3 8B, and Llama 3 70B can be prompted to approximately calibrate their accuracy to specific target percentages on MMLU
IC-101
GPT-4 and Claude 3 struggle to emulate a lower capability profile (high school freshman level) via zero-shot prompting, with only moderate improvement from chain-of-thought prompting