IC-057Command R 7B and 35B rely on each individual pretraining document less per nat of generated information for reasoning than for factual questions, with less volatile influence magnitudes
Laura Ruis, Maximilian Mozes, Juhan Bae, Siddhartha Rao Kamalakara, Dwaraknath Gnaneshwar, Acyr Locatelli, Robert Kirk, Tim Rocktäschel, Edward Grefenstette, Max Bartolo
The total influence per nat of query completion information is consistently lower for reasoning questions than for factual questions at all percentiles of the positive ranking. The variance in influence at the same rank across different queries is also much higher for factual than for reasoning questions, indicating the model relies on more specific, infrequent documents for factual retrieval. The effect is more pronounced for the 35B model. Power-law fits to the top 500 rankings show slightly steeper slopes for reasoning (alpha = -0.36) than factual (alpha = -0.32) for the 35B, meaning a higher percentage of total positive influence is concentrated in the top portions for reasoning.
The authors note the steeper power law for 35B reasoning may be caused by noise, as the query with the steepest slope (alpha = -0.45) has an entirely unrelated top-1 document about lunar eclipses