The paper samples 50 prompts scored above 8 and 50 scored below 2 by GPT-3.5-turbo's difficulty rating, then compares GPT-4 and GPT-3.5-turbo responses on each set. GPT-4 wins 52% of comparisons on the challenging set but only 22% on the easy set, with 40% ties on the easy set. This demonstrates that prompt difficulty is a key factor in distinguishing model capabilities.
Evidence
correlational
Key metric
GPT-4 win rate: 52% on top-50 (score >8) prompts, 22% on bottom-50 (score <2) prompts; 40% tie on bottom-50
Caveat
The comparison is limited to two models (GPT-4 and GPT-3.5-turbo) and only 50 prompts per difficulty tier. The difficulty scores are assigned by GPT-3.5-turbo itself, which may introduce circularity.