IC-208Llama-3-8B-Inst-SimPO does not outperform Llama-3-70B-Inst on WildBench, contrary to its advantage on AlpacaEval-2.0, but performs comparably on information-seeking and creative tasks
Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras, Yejin Choi
On AlpacaEval-2.0, Llama-3-8B-Inst-SimPO (length-controlled win rate 44.7%) significantly outperforms Llama-3-70B-Inst (34.4%). However, on WildBench, the paper finds that Llama-3-8B-Inst-SimPO is still generally worse than Llama-3-70B-Inst (wb-reward mix: 14 vs 21; wb-score: 45.7 vs 50.2). The gap narrows on information-seeking and creative tasks where the 8B SimPO model performs comparably to the 70B. The authors attribute the AlpacaEval discrepancy to task selection bias and weaker evaluation prompting in that benchmark. Llama-3-8B-Inst-SimPO is nonetheless the best 8B model in the WildBench evaluation and outperforms some larger models.
The comparison is on a single benchmark version; the authors note AlpacaEval's task selection bias may explain the discrepancy, but do not control for all confounds.