AmbigDocs: Reasoning across Documents on Different Entities under the Same Name

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Yoonsang, Ye, Xi, Choi, Eunsol
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916351182897152
author Lee, Yoonsang
Ye, Xi
Choi, Eunsol
author_facet Lee, Yoonsang
Ye, Xi
Choi, Eunsol
contents Different entities with the same name can be difficult to distinguish. Handling confusing entity mentions is a crucial skill for language models (LMs). For example, given the question "Where was Michael Jordan educated?" and a set of documents discussing different people named Michael Jordan, can LMs distinguish entity mentions to generate a cohesive answer to the question? To test this ability, we introduce a new benchmark, AmbigDocs. By leveraging Wikipedia's disambiguation pages, we identify a set of documents, belonging to different entities who share an ambiguous name. From these documents, we generate questions containing an ambiguous name and their corresponding sets of answers. Our analysis reveals that current state-of-the-art models often yield ambiguous answers or incorrectly merge information belonging to different entities. We establish an ontology categorizing four types of incomplete answers and automatic evaluation metrics to identify such categories. We lay the foundation for future work on reasoning across multiple documents with ambiguous entities.
format Preprint
id arxiv_https___arxiv_org_abs_2404_12447
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AmbigDocs: Reasoning across Documents on Different Entities under the Same Name
Lee, Yoonsang
Ye, Xi
Choi, Eunsol
Computation and Language
Different entities with the same name can be difficult to distinguish. Handling confusing entity mentions is a crucial skill for language models (LMs). For example, given the question "Where was Michael Jordan educated?" and a set of documents discussing different people named Michael Jordan, can LMs distinguish entity mentions to generate a cohesive answer to the question? To test this ability, we introduce a new benchmark, AmbigDocs. By leveraging Wikipedia's disambiguation pages, we identify a set of documents, belonging to different entities who share an ambiguous name. From these documents, we generate questions containing an ambiguous name and their corresponding sets of answers. Our analysis reveals that current state-of-the-art models often yield ambiguous answers or incorrectly merge information belonging to different entities. We establish an ontology categorizing four types of incomplete answers and automatic evaluation metrics to identify such categories. We lay the foundation for future work on reasoning across multiple documents with ambiguous entities.
title AmbigDocs: Reasoning across Documents on Different Entities under the Same Name
topic Computation and Language
url https://arxiv.org/abs/2404.12447