IC-264All 18 evaluated LLMs fail to abstain when the provided context lacks the answer, with performance gaps of 13.6% to 68.4% relative to the original context

Yifei Ming, Senthil Purushwalkam, Shrey Pandit, Zixuan Ke, Xuan-Phi Nguyen, Caiming Xiong, Shafiq Joty

SourceFaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"

On the unanswerable context task, the context is modified to remove the evidence supporting the ground-truth answer, and the model is instructed to respond 'unknown' if no information is available. Across all 18 chat models, accuracy drops substantially compared to the original context. The gap ranges from 13.6% to 68.4%. Notably, high original-context performance does not predict unanswerable-context performance: phi-3-medium-128k-instruct scores 75.8% on the original context but only 7.4% on the unanswerable version. Within the same family, larger models do better, e.g. llama-3.1-70b-instruct improves by 10.3% over the 8B variant.

Evidence
correlational
Key metric
performance gap 13.6% to 68.4% across all chat models; phi-3-medium-128k-instruct: 75.8% (original) vs 7.4% (unanswerable); llama-3.1-70b-instruct +10.3% over 8B
Caveat
The unanswerable contexts were generated by an LLM (GPT-4o) modifying original contexts, so the difficulty may depend on how well the modification preserves coherence while removing the answer.
Model
Phi-3 Phi-3-mini-128k-instruct, Phi-3-Medium-128K-Instruct, Phi-3.5 Mini Instruct, Llama 3 8B Instruct, 70B Instruct, Llama 3.1 8B Instruct, 70B Instruct, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 Mistral-7B-Instruct-v0.3, Mistral Nemo Instruct 2407, Gemma 2 Gemma-2-9B-IT, Gemma-2-27B-IT, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo, GPT-4o mini, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4 Turbo, Command R+, Claude 3.5 Sonnet
Concepts
Failure mode
Datasets
SQuAD [source], NewsQA [source], TriviaQA [source], Natural Questions / NaturalQA [source], SearchQA [source], HotpotQA [source], BioASQ [source], DROP [source], RACE [source], TextbookQA [source]
Methods
Chain-of-Thought prompting / Chain-of-Thought (CoT) prompting / CoT prompting / Few-shot Chain of Thought / Few-shot CoT / Chain-of-Thought (CoT-S) / CoT-bag prompting / Few-shot CoT prompting / Rationale prompting / Zero-shot CoT prompting / Wei et al. (2022) chain-of-thought prompting [compared-to]
Related findings
IC-265, IC-266, IC-267
Extraction
automatic-extraction