Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
WildBench
anchor
Findings
IC-206
GPT-4o-0513 achieves the highest wb-reward mix score (35.7) on WildBench, with a clear three-tier structure among 40 evaluated LLMs
[eval]
IC-207
Open LLMs (Llama-3-8B-Inst, Yi-1.5-34B-Chat) show weaker performance on coding and math tasks compared to proprietary models (GPT-4-turbo-0409, Claude 3 Opus) which perform well across all task categories
[eval]
IC-208
Llama-3-8B-Inst-SimPO does not outperform Llama-3-70B-Inst on WildBench, contrary to its advantage on AlpacaEval-2.0, but performs comparably on information-seeking and creative tasks
[eval]
IC-601
Lightweight LLMs exhibit high judgment uncertainty (disagreement ratio exceeding 50% for Qwen2-1.5B) when making repeated binary checklist evaluations, with uncertainty increasing as model size decreases
[context]
IC-601
Lightweight LLMs exhibit high judgment uncertainty (disagreement ratio exceeding 50% for Qwen2-1.5B) when making repeated binary checklist evaluations, with uncertainty increasing as model size decreases
[eval]
IC-602
Lightweight LLMs exhibit positional bias in sequential checklist judgments, with judgment inconsistency increasing as the position of the item in the multi-turn dialogue grows
[eval]