SourceReCogLab: a framework testing relational reasoning & cognitive hypotheses on LLMs
The authors ablate the presence of flavor text (story-like embellishments that increase token count while decreasing information density) on social network reasoning tasks. In the majority of models, adding flavor text obstructs reasoning performance on both the number-of-hops and exact-path questions. GPT-4o is the notable exception, showing consistent performance across all four task variants regardless of flavor text presence or question format. The authors also note that some models (Gemini Pro, Gemini Flash, Mixtral-7x22B) paradoxically perform better on the harder exact-path question than the easier number-of-hops question, possibly due to the higher number of decoded tokens.