Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
SummEval
anchor
Findings
IC-134
Mistral Large V2 as a judge on SummEval coherence systematically avoids extreme ratings (1 and 5), while human annotators assign over 24% of items a median rating of 5
[eval]