Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Harnessing Webpage UIs for Text-Rich Visual Understanding
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-163
LLaVA-1.5, LLaVA-Next, and GPT-4V show near-zero accuracy on GUI grounding benchmarks while achieving 50-85 on general image grounding (RefCOCO+), indicating a failure mode specific to GUI grounding scenarios