Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Beyond correlation: The impact of human uncertainty in measuring the effectiveness of automatic evaluation and LLM-as-a-judge
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-134
Mistral Large V2 as a judge on SummEval coherence systematically avoids extreme ratings (1 and 5), while human annotators assign over 24% of items a median rating of 5