IC-320Llama-2-13b-chat underperforms Llama-2-7b-chat on fine-grained dimension-level evaluation

Kehua Feng, Keyan Ding, Jing Yu, Yiwen Qu, Zhiwen Chen, chengfei lv, Gang Yu, Qiang Zhang, Huajun Chen

SourceSaMer: A Scenario-aware Multi-dimensional Evaluator for Large Language Models

On the MD-Eval dataset (360 human-verified samples across 36 scenarios), the larger Llama-2-13b-chat achieves 48.47% dimension-level accuracy and 53.47% overall accuracy, while the smaller Llama-2-7b-chat achieves 53.13% and 53.58% respectively. The authors explicitly flag this as an 'intriguing observation' and state it suggests that increasing model parameters does not necessarily lead to better fine-grained evaluation capabilities. The overall accuracy is nearly identical between the two sizes, so the degradation is concentrated at the dimension level.

Evidence
correlational
Key metric
dim acc. 48.47% (Llama-2-13b-chat) vs 53.13% (Llama-2-7b-chat); overall acc. 53.47% vs 53.58% on MD-Eval
Caveat
The overall accuracy difference is negligible (53.47% vs 53.58%); the gap is only visible at the dimension level. The MD-Eval set is small (360 samples, 10 per scenario).
Model
Llama 2 / Llama 2 base Llama-2-13B-Chat, Llama 2 7B Chat / Llama-2-chat-7b
Concepts
Scale-dependent behaviour
Related findings
IC-321
Extraction
automatic-extraction