Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
2025-01-22
· ICLR 2025 Spotlight ·
anchor
Findings
IC-206
GPT-4o-0513 achieves the highest wb-reward mix score (35.7) on WildBench, with a clear three-tier structure among 40 evaluated LLMs
IC-207
Open LLMs (Llama-3-8B-Inst, Yi-1.5-34B-Chat) show weaker performance on coding and math tasks compared to proprietary models (GPT-4-turbo-0409, Claude 3 Opus) which perform well across all task categories
IC-208
Llama-3-8B-Inst-SimPO does not outperform Llama-3-70B-Inst on WildBench, contrary to its advantage on AlpacaEval-2.0, but performs comparably on information-seeking and creative tasks