Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Vicuna Bench
anchor
Findings
IC-717
Llama-2-Chat's evaluation capability does not improve monotonically with model size
[eval]
IC-718
GPT-4 achieves 0.882 Pearson correlation with human evaluators on 45 customized score rubrics while GPT-3.5-Turbo achieves only 0.392
[eval]