SourceLearning LLM-as-a-Judge for Preference Alignment
The paper evaluates GPT-4o as a generative judge on three self-built commercial preference datasets (creation, math, code) and four public benchmarks. On the domain-specific datasets, GPT-4o scores below both the scalar model and the authors' CON-J. On public benchmarks, GPT-4o achieves competitive performance (e.g., 95.3 on infinity-preference reasoning, 87.6 on pku-saferlhf) but is surpassed by CON-J on infinity-preference and ultrafeedback. The paper concludes that small models trained on domain-specific data can effectively predict domain-related preferences, outperforming a general-purpose frontier model.