The paper evaluates five editing methods (fine-tuning, KE, MEND, SERAC, and IKE) on three LMMs (BLIP-2, MiniGPT-4, LLaVA-1.5) using the proposed MMKE-Bench. Results show that in-context learning-based IKE is strongest for reliability and generalization, memory-based SERAC excels at locality, and parameter-based KE handles portability best. Visual semantic and user-specific edits are more challenging than visual entity edits, evidenced by lower reliability and portability scores.
Evidence
correlational
Key metric
IKE on LLaVA-1.5 achieves t-rel 63.49, i-rel 59.98, i-gen 59.98 for visual entity editing; SERAC achieves t-loc 99.87 and i-loc 99.26; KE achieves port 48.77. All methods fall below 52% on visual semantic portability.
Caveat
The results are contingent on the specific benchmark design and the choice of models and editing methods; they may not generalize to all LMMs or editing scenarios.