Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-681
GPT-3.5 exhibits positional bias when judging which of two LLM responses is superior
IC-682
GPT-3.5 and GPT-4 achieve F1 scores of 0.5820 and 0.6180 respectively on pairwise response quality evaluation against human annotations
IC-683
LLaMA-7B, LLaMA-30B, Vicuna-7B, and Vicuna-13B achieve low accuracy (12.11% to 42.24%) in zero-shot and few-shot log-likelihood response evaluation