Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
RewardBench
anchor
Findings
IC-159
GPT-4o achieves 55.6% accuracy on creation, 74.8% on math, and 68.1% on code as a preference judge, and is outperformed by domain-specific 7B models on those tasks
[eval]
IC-169
BT-based, DPO-based reward models, and GPT-4 as judge all exhibit significant length bias, with their scores correlating with output length rather than quality
[eval]
IC-391
Llama-3 and Qwen-1.5 models exhibit position bias in LM-as-a-judge, retrieval-augmented QA, and math reasoning, with larger models showing less bias
[eval]