Wiki-TabNER: Integrating Named Entity Recognition into Wikipedia Tables

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Koleva, Aneta, Ringsquandl, Martin, Hatem, Ahmed, Runkler, Thomas, Tresp, Volker
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913815149412352
author Koleva, Aneta
Ringsquandl, Martin
Hatem, Ahmed
Runkler, Thomas
Tresp, Volker
author_facet Koleva, Aneta
Ringsquandl, Martin
Hatem, Ahmed
Runkler, Thomas
Tresp, Volker
contents Interest in solving table interpretation tasks has grown over the years, yet it still relies on existing datasets that may be overly simplified. This is potentially reducing the effectiveness of the dataset for thorough evaluation and failing to accurately represent tables as they appear in the real-world. To enrich the existing benchmark datasets, we extract and annotate a new, more challenging dataset. The proposed Wiki-TabNER dataset features complex tables containing several entities per cell, with named entities labeled using DBpedia classes. This dataset is specifically designed to address named entity recognition (NER) task within tables, but it can also be used as a more challenging dataset for evaluating the entity linking task. In this paper we describe the distinguishing features of the Wiki-TabNER dataset and the labeling process. In addition, we propose a prompting framework for evaluating the new large language models on the within tables NER task. Finally, we perform qualitative analysis to gain insights into the challenges encountered by the models and to understand the limitations of the proposed~dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2403_04577
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Wiki-TabNER: Integrating Named Entity Recognition into Wikipedia Tables
Koleva, Aneta
Ringsquandl, Martin
Hatem, Ahmed
Runkler, Thomas
Tresp, Volker
Artificial Intelligence
Computation and Language
Interest in solving table interpretation tasks has grown over the years, yet it still relies on existing datasets that may be overly simplified. This is potentially reducing the effectiveness of the dataset for thorough evaluation and failing to accurately represent tables as they appear in the real-world. To enrich the existing benchmark datasets, we extract and annotate a new, more challenging dataset. The proposed Wiki-TabNER dataset features complex tables containing several entities per cell, with named entities labeled using DBpedia classes. This dataset is specifically designed to address named entity recognition (NER) task within tables, but it can also be used as a more challenging dataset for evaluating the entity linking task. In this paper we describe the distinguishing features of the Wiki-TabNER dataset and the labeling process. In addition, we propose a prompting framework for evaluating the new large language models on the within tables NER task. Finally, we perform qualitative analysis to gain insights into the challenges encountered by the models and to understand the limitations of the proposed~dataset.
title Wiki-TabNER: Integrating Named Entity Recognition into Wikipedia Tables
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2403.04577