UMIE: Unified Multimodal Information Extraction with Instruction Tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Lin, Zhang, Kai, Li, Qingyuan, Lou, Renze
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913187712991232
author Sun, Lin
Zhang, Kai
Li, Qingyuan
Lou, Renze
author_facet Sun, Lin
Zhang, Kai
Li, Qingyuan
Lou, Renze
contents Multimodal information extraction (MIE) gains significant attention as the popularity of multimedia content increases. However, current MIE methods often resort to using task-specific model structures, which results in limited generalizability across tasks and underutilizes shared knowledge across MIE tasks. To address these issues, we propose UMIE, a unified multimodal information extractor to unify three MIE tasks as a generation problem using instruction tuning, being able to effectively extract both textual and visual mentions. Extensive experiments show that our single UMIE outperforms various state-of-the-art (SoTA) methods across six MIE datasets on three tasks. Furthermore, in-depth analysis demonstrates UMIE's strong generalization in the zero-shot setting, robustness to instruction variants, and interpretability. Our research serves as an initial step towards a unified MIE model and initiates the exploration into both instruction tuning and large language models within the MIE domain. Our code, data, and model are available at https://github.com/ZUCC-AI/UMIE
format Preprint
id arxiv_https___arxiv_org_abs_2401_03082
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UMIE: Unified Multimodal Information Extraction with Instruction Tuning
Sun, Lin
Zhang, Kai
Li, Qingyuan
Lou, Renze
Artificial Intelligence
Multimodal information extraction (MIE) gains significant attention as the popularity of multimedia content increases. However, current MIE methods often resort to using task-specific model structures, which results in limited generalizability across tasks and underutilizes shared knowledge across MIE tasks. To address these issues, we propose UMIE, a unified multimodal information extractor to unify three MIE tasks as a generation problem using instruction tuning, being able to effectively extract both textual and visual mentions. Extensive experiments show that our single UMIE outperforms various state-of-the-art (SoTA) methods across six MIE datasets on three tasks. Furthermore, in-depth analysis demonstrates UMIE's strong generalization in the zero-shot setting, robustness to instruction variants, and interpretability. Our research serves as an initial step towards a unified MIE model and initiates the exploration into both instruction tuning and large language models within the MIE domain. Our code, data, and model are available at https://github.com/ZUCC-AI/UMIE
title UMIE: Unified Multimodal Information Extraction with Instruction Tuning
topic Artificial Intelligence
url https://arxiv.org/abs/2401.03082