Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-386
LLM performance on CS-Bench grows logarithmically with parameter scale within model families
IC-387
OpenAI-o1 models substantially improve CS reasoning over GPT-4o at the cost of 14-30x token consumption
IC-388
CS-Bench scores correlate strongly (p > 0.9) with math and code benchmark scores across 12 models
IC-389
All evaluated LLMs score significantly lower on CS reasoning questions than knowledge questions, with the gap narrowing for stronger models