SourceGAIA: a benchmark for General AI Assistants
When evaluating GPT-4 without tools on GAIA level 1 questions, the model achieves non-zero scores on questions that annotators solved via web browsing. The paper attributes these scores primarily to GPT-4 having memorized pieces of information needed to complete intermediate steps, rather than to any actual browsing capability. This means the model's apparent web browsing performance is a shortcut that would fail for questions requiring information not present in its training data. The paper notes that GPT-4 cannot deal with files and multi-modality at all, further confirming its limitations in tool use.