IC-159GPT-4o achieves 55.6% accuracy on creation, 74.8% on math, and 68.1% on code as a preference judge, and is outperformed by domain-specific 7B models on those tasks

Ziyi Ye, Xiangsheng Li, Qiuchi Li, Qingyao Ai, Yujia Zhou, Wei Shen, Dong Yan, Yiqun LIU

SourceLearning LLM-as-a-Judge for Preference Alignment

The paper evaluates GPT-4o as a generative judge on three self-built commercial preference datasets (creation, math, code) and four public benchmarks. On the domain-specific datasets, GPT-4o scores below both the scalar model and the authors' CON-J. On public benchmarks, GPT-4o achieves competitive performance (e.g., 95.3 on infinity-preference reasoning, 87.6 on pku-saferlhf) but is surpassed by CON-J on infinity-preference and ultrafeedback. The paper concludes that small models trained on domain-specific data can effectively predict domain-related preferences, outperforming a general-purpose frontier model.

Evidence
correlational
Key metric
Self-built: creation 55.6, math 74.8, code 68.1. Public: infinity-preference chat 75.0, chat-h 72.2, safety 69.6, reasoning 95.3; ultrafeedback 74.3; pku-saferlhf 87.6; reward-bench 86.9
Model
GPT-4o
Datasets
UltraFeedback [eval], PKU-SafeRLHF [eval], RewardBench [eval]
Extraction
automatic-extraction