IC-587Real-world knowledge acts as a shortcut in LLM relational reasoning, causing worse-than-chance performance on logically valid but factually incongruent statements
Andrew Liu, Henry Prior, Gargi Balasubramaniam, Rivka Moroshko, Amir Zait, Ilia Labzovsky, Danny Karmon, Ishita Dasgupta, Kim Stachenfeld, Kenneth Marino
The authors construct comparison problems using 540 real-world objects with known size/weight, creating congruent (consistent with reality), incongruent (contradicting reality), and random-string (no semantic prior) variants. Across all models, congruent statements outperform incongruent ones. Some models exhibit worse-than-chance (0.5) correctness on incongruent comparison problems, meaning they are actively misled by their real-world priors. Random-string performance lands between the two extremes, confirming that the effect is driven by factual coherence rather than surface form.
Evidence
correlational
Key metric
worse-than-chance (0.5) correctness on incongruent comparison problems; 540 objects used for congruency construction
Caveat
Object sizes and weights were estimated by Gemini Pro, introducing potential noise in the congruency labels.