IC-042No single knowledge editing method excels across all criteria when editing visual and user-specific knowledge in LMMs.

Yuntao Du., Kailin Jiang, Zhi Gao, Chenrui Shi, Zilong Zheng, Siyuan Qi, Qing Li

SourceMMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge

The paper evaluates five editing methods (fine-tuning, KE, MEND, SERAC, and IKE) on three LMMs (BLIP-2, MiniGPT-4, LLaVA-1.5) using the proposed MMKE-Bench. Results show that in-context learning-based IKE is strongest for reliability and generalization, memory-based SERAC excels at locality, and parameter-based KE handles portability best. Visual semantic and user-specific edits are more challenging than visual entity edits, evidenced by lower reliability and portability scores.

Evidence
correlational
Key metric
IKE on LLaVA-1.5 achieves t-rel 63.49, i-rel 59.98, i-gen 59.98 for visual entity editing; SERAC achieves t-loc 99.87 and i-loc 99.26; KE achieves port 48.77. All methods fall below 52% on visual semantic portability.
Caveat
The results are contingent on the specific benchmark design and the choice of models and editing methods; they may not generalize to all LMMs or editing scenarios.
Model
BLIP-2, MiniGPT-4, LLaVA-1.5 / LLaVA-v1.5
Datasets
MMKE-Bench
Methods
Knowledge Editor / KE, MEND, SERAC, In-Context Knowledge Editing / IKE
Extraction
automatic-extraction