SourceTurning large language models into cognitive models
In an appendix accuracy analysis, the paper included GPT-4 as a baseline for modeling human decision-making. GPT-4 achieved 59.72% accuracy on the choices13k dataset and 80.3% on the horizon task. These results were worse than the authors' Centaur model (65.18% and 83.5% respectively) but better than LLaMA without any finetuning.