IC-595Gemini-pro has knowledge gaps on specific topics (Permian extinction, Fordism) causing it to perform far below its average rank on existing benchmarks

Xiang Lisa Li, Farzaan Kaiyom, Evan Zheran Liu, Yifan Mai, Percy Liang, Tatsunori Hashimoto

SourceAutoBencher: Towards Declarative Benchmark Construction

On AutoBench-constructed datasets targeting specific salient topics, Gemini-pro's accuracy drops dramatically compared to its performance on MMLU. On the history dataset (Permian extinction), Gemini-pro scores 0.28 versus 0.84 on MMLU; on the economy dataset (Fordism), it scores 0.48 versus 0.75. Its average rank drops from 6 on existing economic benchmarks to 16 on Fordism. Claude-2.0 also drops by 4 ranks on Permian extinction. In contrast, GPT-3.5-turbo performs better than expected on the 'secret society' topic, rising from rank 7 to 3.

Evidence
correlational
Key metric
Gemini-pro: 0.28 on AutoBench history vs 0.84 on MMLU; 0.48 on AutoBench economy vs 0.75 on MMLU; rank drops from 6 to 16 on Fordism. Claude-2.0 drops 4 ranks on Permian extinction.
Caveat
The specific topics were discovered by an optimization process using GPT-4-turbo as evaluator, which could in principle bias toward topics where GPT-family models are weak. The authors address this by showing Claude-3 models achieve the best accuracies and by a human study confirming the same trends with human-generated questions on the same topics.
Model
Gemini Pro, Claude 2.0, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo-0613
Concepts
Failure mode
Datasets
MMLU / MMLU-Math
Related work
MMLU / MMLU-Math [compared-to]
Related findings
IC-596
Extraction
automatic-extraction