Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
TempCompass
anchor
Findings
IC-510
Qwen2-7B and Llama3-8B score near-random on textual temporal reasoning tasks while Qwen2-72B, Llama3-70B, and GPT-4o achieve near-perfect accuracy, showing temporal reasoning in LLMs is scale-dependent and emerges only above ~70B parameters
[eval]
IC-518
GPT-4o achieves 62.54% overall accuracy on MMWorld, the best among 15 MLLMs, while four open-source models perform below the 26.31% random-choice baseline
[context]