IC-589Flavor text (non-essential descriptive language) degrades relational reasoning in most LLMs, but GPT-4o is robust to it

Andrew Liu, Henry Prior, Gargi Balasubramaniam, Rivka Moroshko, Amir Zait, Ilia Labzovsky, Danny Karmon, Ishita Dasgupta, Kim Stachenfeld, Kenneth Marino

SourceReCogLab: a framework testing relational reasoning & cognitive hypotheses on LLMs

The authors ablate the presence of flavor text (story-like embellishments that increase token count while decreasing information density) on social network reasoning tasks. In the majority of models, adding flavor text obstructs reasoning performance on both the number-of-hops and exact-path questions. GPT-4o is the notable exception, showing consistent performance across all four task variants regardless of flavor text presence or question format. The authors also note that some models (Gemini Pro, Gemini Flash, Mixtral-7x22B) paradoxically perform better on the harder exact-path question than the easier number-of-hops question, possibly due to the higher number of decoded tokens.

Evidence
correlational
Caveat
The effect is described qualitatively from Figure 5; no single unified accuracy delta is reported in the text for the flavor text ablation.
Model
Gemma 2B, Gemma-9B, Gemma-27B, Mixtral 7x22B, Gemini Flash, Pro, GPT-4o
Concepts
Failure mode
Related findings
IC-586, IC-587, IC-588
Extraction
automatic-extraction