IC-871Llama-2-70b-chat as a grader is more generous than GPT-4 and systematically gives higher scores to Llama-2 family outputs

Dingli Yu, Simran Kaur, Arushi Gupta, Jonah Brown-Cohen, Anirudh Goyal, Sanjeev Arora

SourceSKILL-MIX: a Flexible and Expandable Family of Evaluations for AI Models

When both GPT-4 and Llama-2-70b-chat are used to grade the same Skill-Mix generations, Llama-2-70b-chat assigns higher scores overall. More specifically, GPT-4 generally gives a higher score to Mistral-7b-instruct-v0.1 than to Llama-2-7b-chat, but the Llama-2 grader reverses this: it gives a much higher score to Llama-2-7b-chat than to Mistral-7b-instruct-v0.1 for k ≥ 3. The authors conclude via spot-checking that GPT-4 is a more accurate and reliable grader.

Evidence
correlational
Key metric
GPT-4 grading k=3: Mistral-7b-instruct-v0.1 .00/.05/.36 vs Llama-2-7b-chat .00/.01/.25; Llama-2 grading k=3: Mistral-7b-instruct-v0.1 .17/.33/.44 vs Llama-2-7b-chat .33/.50/.70
Caveat
The preference is observed across multiple k values and metrics, but the paper does not quantify the magnitude of the bias in a controlled way beyond the table comparisons.
Model
Llama 2 / Llama 2 base, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 Mistral-7B-Instruct-v0.1
Concepts
Failure mode
Related findings
IC-868, IC-869, IC-870
Extraction
automatic-extraction