Uncovering Entity Identity Confusion in Multimodal Knowledge Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Shu, Ye, Xiaotian, Mou, Xinyu, Liu, Dongsheng, Wang, Xiaohan, Zhang, Mengqi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917467903754240
author Wu, Shu
Ye, Xiaotian
Mou, Xinyu
Liu, Dongsheng
Wang, Xiaohan
Zhang, Mengqi
author_facet Wu, Shu
Ye, Xiaotian
Mou, Xinyu
Liu, Dongsheng
Wang, Xiaohan
Zhang, Mengqi
contents Multimodal knowledge editing (MKE) aims to correct the internal knowledge of large vision-language models after deployment, yet the behavioral patterns of post-edit models remain underexplored. In this paper, we identify a systemic failure mode in edited models, termed Entity Identity Confusion (EIC): edited models exhibit an absurd behavior where text-only queries about the original entity's identity unexpectedly return information about the new entity. To rigorously investigate EIC, we construct EC-Bench, a diagnostic benchmark that directly probes how image-entity bindings shift before and after editing. Our analysis reveals that EIC stems from existing methods failing to distinguish between Image-Entity (I-E) binding and Entity-Entity (E-E) relational knowledge in the model, causing models to overfit E-E associations as a shortcut: the image is still perceived as the original entity, with the new entity's name serving only as a spurious identity label. We further explore potential mitigation strategies, showing that constraining edits to the model's I-E processing stage encourages edits to act more faithfully on I-E binding, thereby substantially reducing EIC. Based on these findings, we discuss principled desiderata for faithful MKE and provide methodological guidance for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06096
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Uncovering Entity Identity Confusion in Multimodal Knowledge Editing
Wu, Shu
Ye, Xiaotian
Mou, Xinyu
Liu, Dongsheng
Wang, Xiaohan
Zhang, Mengqi
Computation and Language
Computer Vision and Pattern Recognition
Multimodal knowledge editing (MKE) aims to correct the internal knowledge of large vision-language models after deployment, yet the behavioral patterns of post-edit models remain underexplored. In this paper, we identify a systemic failure mode in edited models, termed Entity Identity Confusion (EIC): edited models exhibit an absurd behavior where text-only queries about the original entity's identity unexpectedly return information about the new entity. To rigorously investigate EIC, we construct EC-Bench, a diagnostic benchmark that directly probes how image-entity bindings shift before and after editing. Our analysis reveals that EIC stems from existing methods failing to distinguish between Image-Entity (I-E) binding and Entity-Entity (E-E) relational knowledge in the model, causing models to overfit E-E associations as a shortcut: the image is still perceived as the original entity, with the new entity's name serving only as a spurious identity label. We further explore potential mitigation strategies, showing that constraining edits to the model's I-E processing stage encourages edits to act more faithfully on I-E binding, thereby substantially reducing EIC. Based on these findings, we discuss principled desiderata for faithful MKE and provide methodological guidance for future research.
title Uncovering Entity Identity Confusion in Multimodal Knowledge Editing
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.06096