Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
The Alignment Problem from a Deep Learning Perspective
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1240
GPT-4-0314 achieves 85% zero-shot accuracy on situational-awareness questions about its own architecture and training
IC-1241
GPT-4 (14 March 2023) achieves 100% zero-shot accuracy at classifying whether news articles could be part of its pre-training data
IC-1242
GPT-4 wrote a working script that called an instance of itself on its API as part of a plan to gain internet access