When both GPT-4 and Llama-2-70b-chat are used to grade the same Skill-Mix generations, Llama-2-70b-chat assigns higher scores overall. More specifically, GPT-4 generally gives a higher score to Mistral-7b-instruct-v0.1 than to Llama-2-7b-chat, but the Llama-2 grader reverses this: it gives a much higher score to Llama-2-7b-chat than to Mistral-7b-instruct-v0.1 for k ≥ 3. The authors conclude via spot-checking that GPT-4 is a more accurate and reliable grader.
Evidence
correlational
Key metric
GPT-4 grading k=3: Mistral-7b-instruct-v0.1 .00/.05/.36 vs Llama-2-7b-chat .00/.01/.25; Llama-2 grading k=3: Mistral-7b-instruct-v0.1 .17/.33/.44 vs Llama-2-7b-chat .33/.50/.70
Caveat
The preference is observed across multiple k values and metrics, but the paper does not quantify the magnitude of the bias in a controlled way beyond the table comparisons.