The paper measures Pearson correlation between CS-Bench overall scores and scores on math benchmarks (GSM8K, MATH) and code benchmarks (HumanEval, MBPP) for 12 models ranging from Qwen1.5-0.5b to GPT-4. All four math/code correlations exceed 0.9 (GSM8K p=0.93, MATH p=0.94, HumanEval p=0.91, MBPP p=0.96), while correlations with non-math/code benchmarks (GPQA, C-Eval, Tool-Use-Eval) are lower (p < 0.8). This indicates that CS proficiency in LLMs is closely tied to mathematical and programming abilities rather than being an independent capability.
Correlation is measured across only 12 models, limiting statistical power. The paper notes that some models (e.g., Qwen1.5-7b vs Llama2-70b) show inconsistent patterns between math/code and CS, suggesting the correlation is not absolute.