AMELI: Enhancing Multimodal Entity Linking with Fine-Grained Attributes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Barry Menglong, Wang, Sijia, Chen, Yu, Wang, Qifan, Liu, Minqian, Xu, Zhiyang, Yu, Licheng, Huang, Lifu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912422803013632
author Yao, Barry Menglong
Wang, Sijia
Chen, Yu
Wang, Qifan
Liu, Minqian
Xu, Zhiyang
Yu, Licheng
Huang, Lifu
author_facet Yao, Barry Menglong
Wang, Sijia
Chen, Yu
Wang, Qifan
Liu, Minqian
Xu, Zhiyang
Yu, Licheng
Huang, Lifu
contents We propose attribute-aware multimodal entity linking, where the input consists of a mention described with a text paragraph and images, and the goal is to predict the corresponding target entity from a multimodal knowledge base (KB) where each entity is also accompanied by a text description, visual images, and a collection of attributes that present the meta-information of the entity in a structured format. To facilitate this research endeavor, we construct AMELI, encompassing a new multimodal entity linking benchmark dataset that contains 16,735 mentions described in text and associated with 30,472 images, and a multimodal knowledge base that covers 34,690 entities along with 177,873 entity images and 798,216 attributes. To establish baseline performance on AMELI, we experiment with several state-of-the-art architectures for multimodal entity linking and further propose a new approach that incorporates attributes of entities into disambiguation. Experimental results and extensive qualitative analysis demonstrate that extracting and understanding the attributes of mentions from their text descriptions and visual images play a vital role in multimodal entity linking. To the best of our knowledge, we are the first to integrate attributes in the multimodal entity linking task. The programs, model checkpoints, and the dataset are publicly available at https://github.com/VT-NLP/Ameli.
format Preprint
id arxiv_https___arxiv_org_abs_2305_14725
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle AMELI: Enhancing Multimodal Entity Linking with Fine-Grained Attributes
Yao, Barry Menglong
Wang, Sijia
Chen, Yu
Wang, Qifan
Liu, Minqian
Xu, Zhiyang
Yu, Licheng
Huang, Lifu
Computation and Language
I.2.7
We propose attribute-aware multimodal entity linking, where the input consists of a mention described with a text paragraph and images, and the goal is to predict the corresponding target entity from a multimodal knowledge base (KB) where each entity is also accompanied by a text description, visual images, and a collection of attributes that present the meta-information of the entity in a structured format. To facilitate this research endeavor, we construct AMELI, encompassing a new multimodal entity linking benchmark dataset that contains 16,735 mentions described in text and associated with 30,472 images, and a multimodal knowledge base that covers 34,690 entities along with 177,873 entity images and 798,216 attributes. To establish baseline performance on AMELI, we experiment with several state-of-the-art architectures for multimodal entity linking and further propose a new approach that incorporates attributes of entities into disambiguation. Experimental results and extensive qualitative analysis demonstrate that extracting and understanding the attributes of mentions from their text descriptions and visual images play a vital role in multimodal entity linking. To the best of our knowledge, we are the first to integrate attributes in the multimodal entity linking task. The programs, model checkpoints, and the dataset are publicly available at https://github.com/VT-NLP/Ameli.
title AMELI: Enhancing Multimodal Entity Linking with Fine-Grained Attributes
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2305.14725