Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
CREPE
anchor
Findings
IC-1407
OpenFlamingo and Idefics models perform near random chance on compositional image-text matching, and ICL has almost no effect on atomic foils
[eval]
IC-698
GPT-3.5-turbo's CoT reasoning errors are correlated across different demonstration sets, while PoT errors are less correlated
[eval]
IC-699
Llama2-13b produces significantly less consistent answers than GPT-3.5-turbo on complex reasoning tasks, making it unsuitable as a weaker LLM in a cascade
[eval]
IC-700
GPT-4's reasoning accuracy degrades when provided with incorrect hints from a weaker model
[eval]