IC-101GPT-4 and Claude 3 struggle to emulate a lower capability profile (high school freshman level) via zero-shot prompting, with only moderate improvement from chain-of-thought prompting

Teun van der Weij, Felix Hofstätter, Oliver Jaffe, Samuel F. Brown, Francis Rhys Ward

SourceAI Sandbagging: Language Models can Strategically Underperform on Evaluations

When prompted to answer only questions appropriate for a high school freshman and to give wrong answers on harder questions, GPT-4 and Claude 3 mostly fail in the zero-shot setting. GPT-4 only slightly decreases its performance on college-level questions, while Claude 3 moderately decreases performance on both high-school and college questions. With chain-of-thought prompting, performance improves somewhat, but GPT-4 tends to overshoot on college-level questions and Claude 3 undershoots on high-school level questions.

Evidence
correlational
Caveat
The authors note this is a negative result in the zero-shot setting and that the models' inability may be due to difficulty understanding the task rather than a fundamental limitation. Only two models were tested for this specific emulation task.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Claude 3
Concepts
Failure mode
Datasets
MMLU / MMLU-Math [eval]
Methods
Chain-of-Thought prompting / Chain-of-Thought (CoT) prompting / CoT prompting / Few-shot Chain of Thought / Few-shot CoT / Chain-of-Thought (CoT-S) / CoT-bag prompting / Few-shot CoT prompting / Rationale prompting / Zero-shot CoT prompting / Wei et al. (2022) chain-of-thought prompting [compared-to]
Related findings
IC-099, IC-100
Extraction
automatic-extraction