IC-1218GPT-4 achieves 59.72% accuracy on choices13k and 80.3% on the horizon task when modeling human decisions

Marcel Binz, Eric Schulz

SourceTurning large language models into cognitive models

In an appendix accuracy analysis, the paper included GPT-4 as a baseline for modeling human decision-making. GPT-4 achieved 59.72% accuracy on the choices13k dataset and 80.3% on the horizon task. These results were worse than the authors' Centaur model (65.18% and 83.5% respectively) but better than LLaMA without any finetuning.

Evidence
correlational
Key metric
accuracy 59.72% (choices13k), 80.3% (horizon task)
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
Datasets
Choices13k [eval]
Related findings
IC-1217
Extraction
automatic-extraction