IC-410Knowledge editing can degrade generalization performance below pre-edit levels in Llama2-7B, Llama3-8B, and Mistral-v0.3-7B

Baixiang Huang, Canyu Chen, Xiongxiao Xu, Ali Payani, Kai Shu

SourceCan Knowledge Editing Really Correct Hallucinations?

After applying knowledge editing, the models' ability to answer different formulations of the same question (rephrased, yes/no, multi-choice, reversed) can actually drop below the pre-edit level. GRACE degrades generalization across all question types, and FT-L and LoRA degrade on some types. Higher efficacy scores do not correlate with higher generalization scores: GRACE nearly tops efficacy but largely degrades generalization. All editing methods except ICE only slightly improve or negatively impact generalization.

Evidence
interventional
Key metric
post-edit generalization scores could even be lower than pre-edit scores for the same llm and question type; grace degrades across all question types; all editing methods except ice only slightly improve or negatively impact generalization
Caveat
The manifestation of hallucination depends on question design; pre-edit generalization scores are not 0% for each question type.
Model
Llama 2 / Llama 2 base Llama 2 7B, Llama 3 8B, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 Mistral-v0.3 7B
Concepts
Failure mode
Methods
ROME [primary], MEMIT [primary], LoRA [primary], ICE [primary], GRACE [primary]
Related findings
IC-409, IC-411, IC-412
Extraction
automatic-extraction