IC-387OpenAI-o1 models substantially improve CS reasoning over GPT-4o at the cost of 14-30x token consumption

Xiaoshuai Song, Muxi Diao, Guanting Dong, Zhengyang Wang, Yujia Fu, Runqi Qiao, Zhexu Wang, Dayuan Fu, Huangxuan Wu, Bin Liang, Weihao Zeng, Yejie Wang, Zhuoma GongQue, Jianing Yu, Qiuna Tan, Weiran Xu

SourceCS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery

The paper compares OpenAI-o1-mini and o1-preview against GPT-4o on CS-Bench. On reasoning-type questions, o1-mini scores 76.12% (+11.97 over GPT-4o's 64.15%) and o1-preview scores 80.98% (+16.83). On knowledge-type questions the gains are smaller (+0.65 and +6.66 respectively). The performance gain comes at a high token cost: o1-mini consumes 341.8 reasoning tokens (×14.2 over GPT-4o) and o1-preview consumes 732.79 (×30.04). Performance increases roughly logarithmically with token expenditure on reasoning questions.

Evidence
correlational
Key metric
Reasoning: GPT-4o 64.15%, o1-mini 76.12% (+11.97), o1-preview 80.98% (+16.83); Knowledge: GPT-4o 76.95%, o1-mini 77.60% (+0.65), o1-preview 83.61% (+6.66); Overall: GPT-4o 72.29%, o1-mini 77.06% (+4.77), o1-preview 82.65% (+10.36); Tokens: o1-mini ×14.2, o1-preview ×30.04
Caveat
The comparison is limited to English CS-Bench. The token cost makes o1 impractical for many deployment scenarios. The paper notes o1's unique reasoning form required separate analysis from other models.
Model
GPT-4o
Datasets
CS-Bench [eval]
Related findings
IC-386, IC-388, IC-389
Extraction
automatic-extraction