SourceYour Weak LLM is Secretly a Strong Teacher for Alignment
The paper measures GPT-4's consistency as a preference judge by prompting it 10 consecutive times on 100 randomly sampled examples from two sets: dmatch (where a weak LLM and human annotators agree on preference) and dmismatch (where they disagree). In the match set, where quality differences between responses are clearer, GPT-4's majority-vote consistency is 0.84. In the mismatch set, where quality differences are subtle, consistency drops to 0.66. The authors conclude that advanced LLMs like GPT-4 face challenges in providing consistent feedback when response quality distinctions are minimal.