Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Choices13k
anchor
Findings
IC-1217
LLaMA 65B's token-probability readout fails to capture human decision-making, producing near-chance NLL and no human-like exploration behavior
[eval]
IC-1218
GPT-4 achieves 59.72% accuracy on choices13k and 80.3% on the horizon task when modeling human decisions
[eval]
IC-362
LLMs with chain-of-thought prompting predict and simulate human risky choices that are more rational than actual human behavior, correlating more highly with maximum expected value than with human choices
[eval]