M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gan, Chengguang, Cai, Zhixi, Wei, Yanbin, Liang, Yunhao, Ni, Shiwen, Mori, Tatsunori
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909646222000128
author Gan, Chengguang
Cai, Zhixi
Wei, Yanbin
Liang, Yunhao
Ni, Shiwen
Mori, Tatsunori
author_facet Gan, Chengguang
Cai, Zhixi
Wei, Yanbin
Liang, Yunhao
Ni, Shiwen
Mori, Tatsunori
contents Mutual Reinforcement Effect (MRE) is an emerging subfield at the intersection of information extraction and model interpretability. MRE aims to leverage the mutual understanding between tasks of different granularities, enhancing the performance of both coarse-grained and fine-grained tasks through joint modeling. While MRE has been explored and validated in the textual domain, its applicability to visual and multimodal domains remains unexplored. In this work, we extend MRE to the multimodal information extraction domain for the first time. Specifically, we introduce a new task: Multimodal Mutual Reinforcement Effect (M-MRE), and construct a corresponding dataset to support this task. To address the challenges posed by M-MRE, we further propose a Prompt Format Adapter (PFA) that is fully compatible with various Large Vision-Language Models (LVLMs). Experimental results demonstrate that MRE can also be observed in the M-MRE task, a multimodal text-image understanding scenario. This provides strong evidence that MRE facilitates mutual gains across three interrelated tasks, confirming its generalizability beyond the textual domain.
format Preprint
id arxiv_https___arxiv_org_abs_2504_17353
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
Gan, Chengguang
Cai, Zhixi
Wei, Yanbin
Liang, Yunhao
Ni, Shiwen
Mori, Tatsunori
Computation and Language
Computer Vision and Pattern Recognition
Multimedia
Mutual Reinforcement Effect (MRE) is an emerging subfield at the intersection of information extraction and model interpretability. MRE aims to leverage the mutual understanding between tasks of different granularities, enhancing the performance of both coarse-grained and fine-grained tasks through joint modeling. While MRE has been explored and validated in the textual domain, its applicability to visual and multimodal domains remains unexplored. In this work, we extend MRE to the multimodal information extraction domain for the first time. Specifically, we introduce a new task: Multimodal Mutual Reinforcement Effect (M-MRE), and construct a corresponding dataset to support this task. To address the challenges posed by M-MRE, we further propose a Prompt Format Adapter (PFA) that is fully compatible with various Large Vision-Language Models (LVLMs). Experimental results demonstrate that MRE can also be observed in the M-MRE task, a multimodal text-image understanding scenario. This provides strong evidence that MRE facilitates mutual gains across three interrelated tasks, confirming its generalizability beyond the textual domain.
title M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
topic Computation and Language
Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2504.17353