Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Claude Instant 1
anchor
Findings
IC-1330
All 20 evaluated LLMs improve in multi-turn task-solving with additional tool-use turns and GPT-4-simulated language feedback
IC-1361
GPT-4 and other state-of-the-art LLMs achieve near-human accuracy in inferring personal attributes from unstructured text
IC-751
Arena-Hard-200 reveals larger performance gaps between open and proprietary LLMs than MT-Bench