In the interrogative evaluation, GPT-4 and GPT-3.5 first generate content (stories, summaries, dialogue continuations) and are then asked multiple-choice questions about that same generated content. Despite excelling at generation, the models make frequent errors in answering these questions, while humans consistently achieve higher accuracy. The paper notes that humans are assumed to answer questions about their own generations with near-perfect accuracy, so the true gap is likely larger. Qualitative examples show GPT-4 contradicting itself and misreading details of its own stories.
Evidence
correlational
Caveat
The human baseline assumes near-perfect accuracy on questions about one's own generations rather than directly measuring it. The humans in the study are not experts; the authors anticipate the gap would widen with human experts. The study focuses on a small set of the most popular models.