Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Prometheus: Inducing Fine-Grained Evaluation Capability in Language Models
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-717
Llama-2-Chat's evaluation capability does not improve monotonically with model size
IC-718
GPT-4 achieves 0.882 Pearson correlation with human evaluators on 45 customized score rubrics while GPT-3.5-Turbo achieves only 0.392
IC-719
Llama-2-Chat achieves reasonable human-preference accuracy (51.78-53.67%) as a prompted reward model without specific reward-model training