IC-764GPT-4 and GPT-3.5 make frequent errors answering questions about their own generated text, underperforming humans in interrogative evaluation

Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman, Linjie Li, Jena D. Hwang, Liwei Jiang, Jillian Fisher, Abhilasha Ravichander, Khyathi Chandu, Benjamin Newman, Pang Wei Koh, Allyson Ettinger, Yejin Choi

SourceThe Generative AI Paradox: “What It Can Create, It May Not Understand”

In the interrogative evaluation, GPT-4 and GPT-3.5 first generate content (stories, summaries, dialogue continuations) and are then asked multiple-choice questions about that same generated content. Despite excelling at generation, the models make frequent errors in answering these questions, while humans consistently achieve higher accuracy. The paper notes that humans are assumed to answer questions about their own generations with near-perfect accuracy, so the true gap is likely larger. Qualitative examples show GPT-4 contradicting itself and misreading details of its own stories.

Evidence
correlational
Caveat
The human baseline assumes near-perfect accuracy on questions about one's own generations rather than directly measuring it. The humans in the study are not experts; the authors anticipate the gap would widen with human experts. The study focuses on a small set of the most popular models.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5
Concepts
Failure mode
Datasets
XSum [eval], MUTUAL [eval], HellaSwag [eval], COLIE [eval]
Methods
Zero-shot prompting [primary]
Related work
Madaan et al. 2023 (Self-Refine) [context], Agrawal et al. 2023 [context]
Related findings
IC-762, IC-763, IC-765
Extraction
automatic-extraction