Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-056
Influence scores of pretraining documents correlate across reasoning queries of the same type, indicating Command R 7B and 35B rely on shared procedural knowledge rather than retrieving specific answers
IC-057
Command R 7B and 35B rely on each individual pretraining document less per nat of generated information for reasoning than for factual questions, with less volatile influence magnitudes
IC-058
The answer to factual questions appears in the top 0.01% most influential pretraining documents for 55% of 7B queries and 30% of 35B queries, but almost never for reasoning questions
IC-059
Code data is strongly overrepresented in the most influential pretraining documents for reasoning queries in Command R 7B and 35B