SourceThe Trickle-down Impact of Reward Inconsistency on RLHF
The paper evaluates GPT-4 on the contrast instructions benchmark by prompting it to determine which of two instruction-response pairs would be more preferred by humans. Across the four datasets (Stack, WMT, Twitter, RealSumm), GPT-4 achieves 95.5% average response consistency (cres) and 96.0% average instruction consistency (cins). This substantially exceeds the human performance without tools (82.8% cres, 81.7% cins) and is comparable to human performance with tools (96.3% cres, 96.1% cins). The result validates that contrast instructions is a solvable task and that the near-random performance of standard RMs reflects a genuine limitation of the training procedure rather than task impossibility.