IC-388CS-Bench scores correlate strongly (p > 0.9) with math and code benchmark scores across 12 models

Xiaoshuai Song, Muxi Diao, Guanting Dong, Zhengyang Wang, Yujia Fu, Runqi Qiao, Zhexu Wang, Dayuan Fu, Huangxuan Wu, Bin Liang, Weihao Zeng, Yejie Wang, Zhuoma GongQue, Jianing Yu, Qiuna Tan, Weiran Xu

SourceCS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery

The paper measures Pearson correlation between CS-Bench overall scores and scores on math benchmarks (GSM8K, MATH) and code benchmarks (HumanEval, MBPP) for 12 models ranging from Qwen1.5-0.5b to GPT-4. All four math/code correlations exceed 0.9 (GSM8K p=0.93, MATH p=0.94, HumanEval p=0.91, MBPP p=0.96), while correlations with non-math/code benchmarks (GPQA, C-Eval, Tool-Use-Eval) are lower (p < 0.8). This indicates that CS proficiency in LLMs is closely tied to mathematical and programming abilities rather than being an independent capability.

Evidence
correlational
Key metric
Pearson p: GSM8K 0.93, MATH 0.94, HumanEval 0.91, MBPP 0.96; non-CS benchmarks: GPQA 0.7175, C-Eval 0.7983, Tool-Use-Eval 0.4019, T-Eval 0.5282
Caveat
Correlation is measured across only 12 models, limiting statistical power. The paper notes that some models (e.g., Qwen1.5-7b vs Llama2-70b) show inconsistent patterns between math/code and CS, suggesting the correlation is not absolute.
Model
Qwen1.5, Llama 2 / Llama 2 base, Mixtral 8x7B / Mistral 8x7B Instruct / Mixtral 46.7B / Mixtral 8x7B Instruct / Mixtral-instruct-8x7b, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
Datasets
CS-Bench [eval], GSM8K [eval], HumanEval [eval], MBPP [eval], GPQA [eval], C-Eval [eval], Chatbot Arena [eval]
Methods
Pearson correlation coefficient [primary]
Related findings
IC-386, IC-387, IC-389
Extraction
automatic-extraction