IC-209LLM judges (GPT-3.5-turbo-1106, GPT-4o-mini, GPT-4o, Claude-3-5-sonnet) implicitly prioritize style over factuality and safety when scoring pairwise preferences

Benjamin Feuer, Micah Goldblum, Teresa Datta, Sanjana Nambiar, Raz Besaleli, Samuel Dooley, Max Cembalest, John P Dickerson

SourceStyle Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking

When prompted with explicit criteria (completeness, conciseness, style, safety, correctness), all four judge models in the panel assign a style score that correlates almost perfectly with their overall preference judgment (Pearson r = 0.999), while safety (r = 0.727) and conciseness (r = 0.114) are far weaker predictors. When systematic violations are introduced into model responses, the judge penalizes stylistic changes (sarcasm: 96% score loss, conciseness: 63%) far more heavily than factual errors (13%) or blandness (8%). This asymmetry is consistent across all four judges and is not explained by the explicit instructions given to the judge.

Evidence
correlational
Key metric
style Pearson r = 0.999 (std 0), completeness r = 0.963, correctness r = 0.881, safety r = 0.727, conciseness r = 0.114 (Table 2); violation losses: bland 08%, wrong 13%, concise 63%, sarcastic 96% (Table 3)
Caveat
The authors note that factors are not necessarily independent and that the analysis is limited to arena-hard-auto; they also acknowledge that varying components in the llm-judge pipeline could alter behavior substantially, so results should not be treated as evidence that all llm-judge benchmarks follow the same inductive bias.
Model
GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo-1106, GPT-4o GPT-4o-mini-2024-07-18, GPT-4o-2024-08-06
Concepts
Shortcut, Failure mode, Method artefact
Datasets
Arena-Hard-Auto [eval]
Methods
Arena-Hard-Auto [eval]
Related findings
IC-210
Extraction
automatic-extraction