SourceThe False Promise of Imitating Proprietary Language Models
The paper uses GPT-4 to conduct blind pairwise comparisons between ChatGPT outputs and the authors' imitation model outputs, following the same procedure as their Mechanical Turk evaluations. GPT-4's ratings show the same trends as human raters: it rates imitation models as competitive with ChatGPT, particularly when the imitation models produce stylistically similar (confident, well-structured) outputs. The authors note this implies that LLMs may replicate human-like cognitive biases in evaluation, and that GPT-4 may be a viable cheap proxy for human evaluation on some tasks, albeit with a slightly larger absolute preference for ChatGPT's outputs.