IC-751Arena-Hard-200 reveals larger performance gaps between open and proprietary LLMs than MT-Bench

Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric Xing, Joseph E. Gonzalez, Ion Stoica, Hao Zhang

SourceLMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

The paper evaluates ten released models on Arena-Hard-200, a benchmark of 200 challenging user prompts selected from Chatbot Arena conversations. The results show that proprietary models (GPT-4, Claude-2, Claude-instant-v1) score substantially higher than open models (Vicuna, Llama-2, WizardLM), and the gap is wider than what MT-Bench shows. The paper concludes this suggests more room for open models to catch up on challenging real-world tasks.

Evidence
correlational
Caveat
Scores are assigned by GPT-4 as judge, introducing potential bias in favor of GPT-4's own style. The paper does not print exact per-model scores in the text; they appear only in Figure 6.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo, Claude Instant 1, Vicuna Vicuna-7B-v1.5, Vicuna-13b-v1.5, Vicuna-33B-v1.3, Llama 2 / Llama 2 base Llama 2 7B Chat / Llama-2-chat-7b, Llama-2-70B-Chat, Wizardlm WizardLM-70B-v1.0
Datasets
Arena-Hard-200 [eval], MT-Bench
Methods
LLM-as-a-Judge / GPT-4 as judge / GPT-4o as LLM judge [eval]
Related work
MT-Bench [compared-to]
Related findings
IC-750, IC-752
Extraction
automatic-extraction