Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
G-Eval
anchor
Findings
IC-312
GPT-4o achieves the highest insight-level Llama-3-eval score (0.60) among all tested LLM backbones on InsightBench multi-step data analytics
[context]
IC-314
Llama-3-70b as an LLM-based evaluator (Llama-3-eval) produces agent rankings consistent with GPT-4-based G-Eval on InsightBench
[builds-on]
IC-314
Llama-3-70b as an LLM-based evaluator (Llama-3-eval) produces agent rankings consistent with GPT-4-based G-Eval on InsightBench
[compared-to]