Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Kendall's tau / Kendall-tau
Findings
IC-095
Evaluated LLMs achieve at most 36% latest-state accuracy and 18% full-set accuracy on multi-event entity tracking, with low Kendall's tau on chronological ordering
[eval]
IC-717
Llama-2-Chat's evaluation capability does not improve monotonically with model size
[eval]