Enhancing Multimodal Entity and Relation Extraction with Variational Information Bottleneck

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cui, Shiyao, Cao, Jiangxia, Cong, Xin, Sheng, Jiawei, Li, Quangang, Liu, Tingwen, Shi, Jinqiao
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909099179900928
author Cui, Shiyao
Cao, Jiangxia
Cong, Xin
Sheng, Jiawei
Li, Quangang
Liu, Tingwen
Shi, Jinqiao
author_facet Cui, Shiyao
Cao, Jiangxia
Cong, Xin
Sheng, Jiawei
Li, Quangang
Liu, Tingwen
Shi, Jinqiao
contents This paper studies the multimodal named entity recognition (MNER) and multimodal relation extraction (MRE), which are important for multimedia social platform analysis. The core of MNER and MRE lies in incorporating evident visual information to enhance textual semantics, where two issues inherently demand investigations. The first issue is modality-noise, where the task-irrelevant information in each modality may be noises misleading the task prediction. The second issue is modality-gap, where representations from different modalities are inconsistent, preventing from building the semantic alignment between the text and image. To address these issues, we propose a novel method for MNER and MRE by Multi-Modal representation learning with Information Bottleneck (MMIB). For the first issue, a refinement-regularizer probes the information-bottleneck principle to balance the predictive evidence and noisy information, yielding expressive representations for prediction. For the second issue, an alignment-regularizer is proposed, where a mutual information-based item works in a contrastive manner to regularize the consistent text-image representations. To our best knowledge, we are the first to explore variational IB estimation for MNER and MRE. Experiments show that MMIB achieves the state-of-the-art performances on three public benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2304_02328
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Enhancing Multimodal Entity and Relation Extraction with Variational Information Bottleneck
Cui, Shiyao
Cao, Jiangxia
Cong, Xin
Sheng, Jiawei
Li, Quangang
Liu, Tingwen
Shi, Jinqiao
Multimedia
Computation and Language
This paper studies the multimodal named entity recognition (MNER) and multimodal relation extraction (MRE), which are important for multimedia social platform analysis. The core of MNER and MRE lies in incorporating evident visual information to enhance textual semantics, where two issues inherently demand investigations. The first issue is modality-noise, where the task-irrelevant information in each modality may be noises misleading the task prediction. The second issue is modality-gap, where representations from different modalities are inconsistent, preventing from building the semantic alignment between the text and image. To address these issues, we propose a novel method for MNER and MRE by Multi-Modal representation learning with Information Bottleneck (MMIB). For the first issue, a refinement-regularizer probes the information-bottleneck principle to balance the predictive evidence and noisy information, yielding expressive representations for prediction. For the second issue, an alignment-regularizer is proposed, where a mutual information-based item works in a contrastive manner to regularize the consistent text-image representations. To our best knowledge, we are the first to explore variational IB estimation for MNER and MRE. Experiments show that MMIB achieves the state-of-the-art performances on three public benchmarks.
title Enhancing Multimodal Entity and Relation Extraction with Variational Information Bottleneck
topic Multimedia
Computation and Language
url https://arxiv.org/abs/2304.02328