IC-266GPT-4o drops from 96.3% closed-book accuracy to 47.5% when given counterfactual context that contradicts its parametric knowledge, far below the 95% human accuracy on the same items

Yifei Ming, Senthil Purushwalkam, Shrey Pandit, Zixuan Ke, Xuan-Phi Nguyen, Caiming Xiong, Shafiq Joty

SourceFaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"

On the counterfactual context task, a multi-paragraph context is generated that provides fabricated evidence supporting an answer that contradicts well-known facts (e.g., 'wood is magnetic'). In the closed-book setting (no context), nearly half the models exceed 90% accuracy, with GPT-4o at 96.3%. When the counterfactual context is added, performance drops sharply: GPT-4o falls to 47.5%. A human study on a held-out subset shows 95% accuracy, indicating the correct answer is derivable from the context. The gap between model and human performance highlights a faithfulness limitation: models cannot override their parametric knowledge in favor of the provided context.

Evidence
correlational
Key metric
GPT-4o: 96.3% (closed-book) vs 47.5% (counterfactual context); human accuracy 95% on held-out subset; nearly half of models achieve over 90% closed-book
Caveat
The counterfactual contexts were generated by GPT-4o and validated via string matching (68.9% pass rate). The task uses multiple-choice format from ARC-Challenge, which may constrain the type of counterfactuals tested.
Model
Phi-3 Phi-3-mini-128k-instruct, Phi-3-Medium-128K-Instruct, Phi-3.5 Mini Instruct, Llama 3 8B Instruct, 70B Instruct, Llama 3.1 8B Instruct, 70B Instruct, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 Mistral-7B-Instruct-v0.3, Mistral Nemo Instruct 2407, Gemma 2 Gemma-2-9B-IT, Gemma-2-27B-IT, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo, GPT-4o mini, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4 Turbo, Command R+, Claude 3.5 Sonnet
Concepts
Failure mode
Datasets
ARC-Challenge [source]
Related findings
IC-264, IC-265, IC-267
Extraction
automatic-extraction