Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
WebArena-Lite
anchor
Findings
IC-477
Released LLMs achieve limited success rates as web agents on WebArena-Lite, with open-source models substantially below proprietary ones
[eval]
IC-478
GPT-4 and GPT-4V achieve approximately 71-73% accuracy in judging whether a web agent trajectory successfully completes a task
[eval]