The paper evaluates ten released models on Arena-Hard-200, a benchmark of 200 challenging user prompts selected from Chatbot Arena conversations. The results show that proprietary models (GPT-4, Claude-2, Claude-instant-v1) score substantially higher than open models (Vicuna, Llama-2, WizardLM), and the gap is wider than what MT-Bench shows. The paper concludes this suggests more room for open models to catch up on challenging real-world tasks.
Evidence
correlational
Caveat
Scores are assigned by GPT-4 as judge, introducing potential bias in favor of GPT-4's own style. The paper does not print exact per-model scores in the text; they appear only in Figure 6.