IC-718GPT-4 achieves 0.882 Pearson correlation with human evaluators on 45 customized score rubrics while GPT-3.5-Turbo achieves only 0.392

Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, Minjoon Seo

SourcePrometheus: Inducing Fine-Grained Evaluation Capability in Language Models

The paper measures how well GPT-4 and GPT-3.5-Turbo can assign scores to long-form responses based on user-defined customized score rubrics, compared against human annotators. On 45 instances each with a unique rubric (drawn from Feedback Bench, MT Bench, and Vicuna Bench), GPT-4 reaches 0.882 Pearson correlation with humans, while GPT-3.5-Turbo reaches only 0.392. In a pairwise feedback-quality comparison, GPT-4 is preferred over GPT-3.5-Turbo, and GPT-4's feedback is characterized as more neutral and abstract.

Evidence
correlational
Key metric
Pearson correlation with human evaluators on 45 customized rubrics: GPT-4 0.882, GPT-3.5-Turbo 0.392
Caveat
The human evaluation set excluded coding and math-related questions; the 45 instances are a small sample. The authors note that GPT-4's self-consistency (measured by sampling 6 times) is lower than its correlation with humans, suggesting some variance in GPT-4's own scoring.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo
Datasets
Vicuna Bench [eval], MT-Bench [eval]
Methods
Pearson correlation [eval]
Related findings
IC-717, IC-719
Extraction
automatic-extraction