IC-1240GPT-4-0314 achieves 85% zero-shot accuracy on situational-awareness questions about its own architecture and training

Richard Ngo, Lawrence Chan, Sören Mindermann

SourceThe Alignment Problem from a Deep Learning Perspective

The authors evaluated GPT-4-0314 on a set of multiple-choice questions probing technical self-knowledge (e.g. parameter counts, training data, attention layer details) drawn from the dataset in Perez et al. (2022b). Using zero-shot prompting with a system message instructing a single-character answer at temperature 0, the model reached 85% accuracy. The authors note this is a preliminary test and that the questions were designed for models similar to Anthropic's, not specifically for GPT-4.

Evidence
correlational
Key metric
85% zero-shot accuracy on situational awareness questions (Appendix E)
Caveat
The dataset was designed for Anthropic models; the authors note results are preliminary and that the questions may not be fully applicable to GPT-4's architecture. No chain-of-thought or other prompting techniques were used.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
Datasets
Fictional Knowledge Dataset / Perez et al. (2022b) self-knowledge dataset [eval]
Related work
Perez et al. (2022b) – Discovering language model behaviors with model-written evaluations [builds-on]
Related findings
IC-1241, IC-1242
Extraction
automatic-extraction