Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
MUTUAL
anchor
Findings
IC-762
GPT-4 and GPT-3.5 outperform humans in generation but underperform in discriminative (selective) evaluation across 10 of 13 language tasks
[eval]
IC-764
GPT-4 and GPT-3.5 make frequent errors answering questions about their own generated text, underperforming humans in interrogative evaluation
[eval]