MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ramamoorthy, Sathyanarayanan, Shah, Vishwa, Khanuja, Simran, Sheikh, Zaid, Jie, Shan, Chia, Ann, Chua, Shearman, Neubig, Graham
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908596863762432
author Ramamoorthy, Sathyanarayanan
Shah, Vishwa
Khanuja, Simran
Sheikh, Zaid
Jie, Shan
Chia, Ann
Chua, Shearman
Neubig, Graham
author_facet Ramamoorthy, Sathyanarayanan
Shah, Vishwa
Khanuja, Simran
Sheikh, Zaid
Jie, Shan
Chia, Ann
Chua, Shearman
Neubig, Graham
contents This paper introduces MERLIN, a novel testbed system for the task of Multilingual Multimodal Entity Linking. The created dataset includes BBC news article titles, paired with corresponding images, in five languages: Hindi, Japanese, Indonesian, Vietnamese, and Tamil, featuring over 7,000 named entity mentions linked to 2,500 unique Wikidata entities. We also include several benchmarks using multilingual and multimodal entity linking methods exploring different language models like LLaMa-2 and Aya-23. Our findings indicate that incorporating visual data improves the accuracy of entity linking, especially for entities where the textual context is ambiguous or insufficient, and particularly for models that do not have strong multilingual abilities. For the work, the dataset, methods are available here at https://github.com/rsathya4802/merlin
format Preprint
id arxiv_https___arxiv_org_abs_2510_14307
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking
Ramamoorthy, Sathyanarayanan
Shah, Vishwa
Khanuja, Simran
Sheikh, Zaid
Jie, Shan
Chia, Ann
Chua, Shearman
Neubig, Graham
Computation and Language
Artificial Intelligence
This paper introduces MERLIN, a novel testbed system for the task of Multilingual Multimodal Entity Linking. The created dataset includes BBC news article titles, paired with corresponding images, in five languages: Hindi, Japanese, Indonesian, Vietnamese, and Tamil, featuring over 7,000 named entity mentions linked to 2,500 unique Wikidata entities. We also include several benchmarks using multilingual and multimodal entity linking methods exploring different language models like LLaMa-2 and Aya-23. Our findings indicate that incorporating visual data improves the accuracy of entity linking, especially for entities where the textual context is ambiguous or insufficient, and particularly for models that do not have strong multilingual abilities. For the work, the dataset, methods are available here at https://github.com/rsathya4802/merlin
title MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.14307