IC-1243GPT-4 and GPT-4 Turbo achieve near-zero scores on GAIA level 3 and single-digit to low-double-digit scores on levels 1-2, compared to 87-94% for human annotators
Grégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun, Thomas Scialom
The paper evaluates GPT-4, GPT-4 Turbo, AutoGPT (GPT-4 backend), and GPT-4 with manually selected plugins on 466 GAIA questions across three difficulty levels. All LLM variants score 0% on level 3. On level 1, GPT-4 scores 9.1%, GPT-4 Turbo 13.0%, AutoGPT 14.4%, and GPT-4+plugins 30.3%. On level 2, scores drop to 2.6%, 5.5%, 0.4%, and 9.7% respectively. Human annotators score 93.9%, 91.8%, and 87.3% across the three levels. The results demonstrate a fundamental failure of current LLMs in multi-step real-world task completion requiring tool use, web browsing, and multi-modality handling.
GPT-4+plugins scores were obtained by manually selecting plugins per question and cannot be reproduced exactly; plugins change or disappear from the store. The paper only evaluates the strongest available LLMs with tool access.