IC-1243GPT-4 and GPT-4 Turbo achieve near-zero scores on GAIA level 3 and single-digit to low-double-digit scores on levels 1-2, compared to 87-94% for human annotators

Grégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun, Thomas Scialom

SourceGAIA: a benchmark for General AI Assistants

The paper evaluates GPT-4, GPT-4 Turbo, AutoGPT (GPT-4 backend), and GPT-4 with manually selected plugins on 466 GAIA questions across three difficulty levels. All LLM variants score 0% on level 3. On level 1, GPT-4 scores 9.1%, GPT-4 Turbo 13.0%, AutoGPT 14.4%, and GPT-4+plugins 30.3%. On level 2, scores drop to 2.6%, 5.5%, 0.4%, and 9.7% respectively. Human annotators score 93.9%, 91.8%, and 87.3% across the three levels. The results demonstrate a fundamental failure of current LLMs in multi-step real-world task completion requiring tool use, web browsing, and multi-modality handling.

Evidence
correlational
Key metric
GPT-4: 9.1 ± 2.5 (L1), 2.6 ± 0.6 (L2), 0 (L3); GPT-4 Turbo: 13.0 ± 2.1 (L1), 5.5 ± 1.4 (L2), 0 (L3); AutoGPT (GPT-4 backend): 14.4 (L1), 0.4 (L2), 0 (L3); GPT-4+plugins: 30.3 (L1), 9.7 (L2), 0 (L3); Human: 93.9 (L1), 91.8 (L2), 87.3 (L3)
Caveat
GPT-4+plugins scores were obtained by manually selecting plugins per question and cannot be reproduced exactly; plugins change or disappear from the store. The paper only evaluates the strongest available LLMs with tool access.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
Concepts
Failure mode
Datasets
GAIA [eval]
Methods
AutoGPT [compared-to]
Related work
MMLU / MMLU-Math [context], GSM8K [context]
Related findings
IC-1244
Extraction
automatic-extraction