Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
LLM-as-a-Judge / GPT-4 as judge / GPT-4o as LLM judge
Findings
IC-121
Gemini 1.0 Pro, when prompted as a zero-shot chain-of-thought judge with majority voting, underperforms fine-tuned smaller Gemma models as verifiers on GSM8K
[compared-to]
IC-748
Pruned LLMs at ≥50% sparsity remain robust in-context retrievers and summarizers, with Vicuna-7B matching up to ~40% sparsity and Vicuna-13B up to ~50% sparsity in open-book settings
[eval]
IC-751
Arena-Hard-200 reveals larger performance gaps between open and proprietary LLMs than MT-Bench
[eval]
IC-811
Alpaca's 52k instruction-tuning data is predominantly low-quality (only 17.75% score ≥ 4.5 on accuracy), yet the full 52k data still yields higher MMLU scores than the filtered 9k subset for both 7b and 13b variants
[eval]