IC-059Code data is strongly overrepresented in the most influential pretraining documents for reasoning queries in Command R 7B and 35B
Laura Ruis, Maximilian Mozes, Juhan Bae, Siddhartha Rao Kamalakara, Dwaraknath Gnaneshwar, Acyr Locatelli, Robert Kirk, Tim Rocktäschel, Edward Grefenstette, Max Bartolo
Comparing the source datasets of the top and bottom k influential documents to the pretraining distribution, code sources are heavily overrepresented for reasoning queries. StackExchange has 10x more influential data than expected for the top portions, and other code sources are 2x overrepresented for k = 50 to 50000. For the 7B, StackExchange and math/trivia have multipliers of 50 and 24 respectively at k=50 for reasoning, versus 27 and 5 for factual. For the 35B, the multipliers are 62 and 21 for reasoning. Content analysis of the top 500 most frequent documents among the k=5000 most influential confirms the 'code' category dominates, while random samples show a much smaller fraction of code.
The authors note this is conventional wisdom among practitioners and cite prior work (Aryabumi et al. 2024) finding code important for reasoning; the question of why code is influential remains open