IC-207Open LLMs (Llama-3-8B-Inst, Yi-1.5-34B-Chat) show weaker performance on coding and math tasks compared to proprietary models (GPT-4-turbo-0409, Claude 3 Opus) which perform well across all task categories
Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras, Yejin Choi
The paper breaks down wb-score performance across five consolidated task categories (reasoning & planning, creative tasks, coding & debugging, info seeking, math & data) for six models. Larger proprietary models like GPT-4-turbo-0409 and Claude 3 Opus perform well across all categories. In contrast, open LLMs like Llama-3-8B-Inst and Yi-1.5-34B-Chat show notably weaker performance specifically on coding and math-related tasks, while their performance on information-seeking and creative tasks is closer to the top models. This reveals a task-specific weakness pattern in open models that is not visible in the aggregate score.
Evidence
correlational
Caveat
The task-category breakdown is shown in a radar plot (Figure 5) for only 6 selected models; exact per-category scores are not printed in the text.