Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Abductive NLI / Abductive NLG
anchor
Note
written α-NLI and α-NLG in the source; see the note on dataset:defeasible-nli
Findings
IC-762
GPT-4 and GPT-3.5 outperform humans in generation but underperform in discriminative (selective) evaluation across 10 of 13 language tasks
[eval]