Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
SaMer: A Scenario-aware Multi-dimensional Evaluator for Large Language Models
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-320
Llama-2-13b-chat underperforms Llama-2-7b-chat on fine-grained dimension-level evaluation
IC-321
GPT-4o selects evaluation dimensions with high precision but low recall, indicating a selective rather than comprehensive strategy