IC-1244GPT-4's non-zero scores on GAIA web browsing questions are largely due to memorization of intermediate information from training data rather than actual web browsing

Grégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun, Thomas Scialom

SourceGAIA: a benchmark for General AI Assistants

When evaluating GPT-4 without tools on GAIA level 1 questions, the model achieves non-zero scores on questions that annotators solved via web browsing. The paper attributes these scores primarily to GPT-4 having memorized pieces of information needed to complete intermediate steps, rather than to any actual browsing capability. This means the model's apparent web browsing performance is a shortcut that would fail for questions requiring information not present in its training data. The paper notes that GPT-4 cannot deal with files and multi-modality at all, further confirming its limitations in tool use.

Evidence
correlational
Caveat
The paper does not provide a precise count of how many web browsing questions were answered via memorization versus other means; the attribution is qualitative ('mostly due to correct memorization').
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
Concepts
Shortcut
Datasets
GAIA [eval]
Related findings
IC-1243
Extraction
automatic-extraction