Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Style Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-209
LLM judges (GPT-3.5-turbo-1106, GPT-4o-mini, GPT-4o, Claude-3-5-sonnet) implicitly prioritize style over factuality and safety when scoring pairwise preferences
IC-210
GPT-4o-mini-2024-07-18 does not exhibit authority bias when used as a judge: appending fabricated references to model responses decreases rather than increases its score