IC-210GPT-4o-mini-2024-07-18 does not exhibit authority bias when used as a judge: appending fabricated references to model responses decreases rather than increases its score

Benjamin Feuer, Micah Goldblum, Teresa Datta, Sanjana Nambiar, Raz Besaleli, Samuel Dooley, Max Cembalest, John P Dickerson

SourceStyle Outweighs Substance: Failure Modes of LLM Judges in Alignment Benchmarking

The paper tests whether adding authoritative-looking references to a model's response can game the LLM judge, a vulnerability reported in prior work. Two hacks were tried: (1) appending generic verified references matched to the question category, and (2) imitating the judge's own response style with 5-7 tailored references in MLA format. Both hacks decreased the arena-hard-auto score by approximately 48-49%, and the judge explicitly flagged the citations as unnecessary. This shows GPT-4o-mini is robust to this particular manipulation when acting as a judge.

Evidence
correlational
Key metric
Hack 1 loss: 49%, Hack 2 loss: 48% (Table 11)
Caveat
The authors note that the authority bias 'remains potentially hackable' but 'must be executed with precision to avoid detection, at least when the judge is a foundation model such as gpt-4o-mini.' Only one judge model was tested for this specific hack.
Model
GPT-4o GPT-4o-mini-2024-07-18
Datasets
Arena-Hard-Auto [eval]
Methods
Arena-Hard-Auto [eval]
Related findings
IC-209
Extraction
automatic-extraction