After applying knowledge editing, the models' ability to answer different formulations of the same question (rephrased, yes/no, multi-choice, reversed) can actually drop below the pre-edit level. GRACE degrades generalization across all question types, and FT-L and LoRA degrade on some types. Higher efficacy scores do not correlate with higher generalization scores: GRACE nearly tops efficacy but largely degrades generalization. All editing methods except ICE only slightly improve or negatively impact generalization.
Evidence
interventional
Key metric
post-edit generalization scores could even be lower than pre-edit scores for the same llm and question type; grace degrades across all question types; all editing methods except ice only slightly improve or negatively impact generalization
Caveat
The manifestation of hallucination depends on question design; pre-edit generalization scores are not 0% for each question type.