IC-752GPT-4's win rate over GPT-3.5-turbo is 52% on the top-50 most challenging prompts but only 22% on the bottom-50 easiest prompts

Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric Xing, Joseph E. Gonzalez, Ion Stoica, Hao Zhang

SourceLMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

The paper samples 50 prompts scored above 8 and 50 scored below 2 by GPT-3.5-turbo's difficulty rating, then compares GPT-4 and GPT-3.5-turbo responses on each set. GPT-4 wins 52% of comparisons on the challenging set but only 22% on the easy set, with 40% ties on the easy set. This demonstrates that prompt difficulty is a key factor in distinguishing model capabilities.

Evidence
correlational
Key metric
GPT-4 win rate: 52% on top-50 (score >8) prompts, 22% on bottom-50 (score <2) prompts; 40% tie on bottom-50
Caveat
The comparison is limited to two models (GPT-4 and GPT-3.5-turbo) and only 50 prompts per difficulty tier. The difficulty scores are assigned by GPT-3.5-turbo itself, which may introduce circularity.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo
Datasets
LMSYS-Chat-1M [source]
Related findings
IC-750, IC-751
Extraction
automatic-extraction