Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Zero-shot prompting
Findings
IC-1584
LLaMA-7B and GPT-J-6B fail to interpret textual emphasis markers, with marked prompting degrading performance substantially
[compared-to]
IC-1585
LLaMA-7B and GPT-J-6B exhibit positional bias in instruction following: zero-shot performance varies significantly when the instruction is moved from after to before the context
[primary]
IC-762
GPT-4 and GPT-3.5 outperform humans in generation but underperform in discriminative (selective) evaluation across 10 of 13 language tasks
[primary]
IC-763
CLIP and OpenCLIP fall short of human discriminative accuracy on vision tasks, with performance dropping substantially under hard negatives
[primary]
IC-764
GPT-4 and GPT-3.5 make frequent errors answering questions about their own generated text, underperforming humans in interrogative evaluation
[primary]
IC-765
BLIP-2, BLIP, InstructBLIP, Bard, and BingChat fall short of human accuracy in answering questions about Midjourney-generated images
[primary]
IC-858
The choice of graph encoding method significantly changes LLM accuracy on graph reasoning tasks, with incident encoding outperforming adjacency by up to 34 percentage points on connected nodes
[primary]