OpenPI2.0: An Improved Dataset for Entity Tracking in Texts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Li, Xu, Hainiu, Kommula, Abhinav, Callison-Burch, Chris, Tandon, Niket
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911764174602240
author Zhang, Li
Xu, Hainiu
Kommula, Abhinav
Callison-Burch, Chris
Tandon, Niket
author_facet Zhang, Li
Xu, Hainiu
Kommula, Abhinav
Callison-Burch, Chris
Tandon, Niket
contents Much text describes a changing world (e.g., procedures, stories, newswires), and understanding them requires tracking how entities change. An earlier dataset, OpenPI, provided crowdsourced annotations of entity state changes in text. However, a major limitation was that those annotations were free-form and did not identify salient changes, hampering model evaluation. To overcome these limitations, we present an improved dataset, OpenPI2.0, where entities and attributes are fully canonicalized and additional entity salience annotations are added. On our fairer evaluation setting, we find that current state-of-the-art language models are far from competent. We also show that using state changes of salient entities as a chain-of-thought prompt, downstream performance is improved on tasks such as question answering and classical planning, outperforming the setting involving all related entities indiscriminately. We offer OpenPI2.0 for the continued development of models that can understand the dynamics of entities in text.
format Preprint
id arxiv_https___arxiv_org_abs_2305_14603
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle OpenPI2.0: An Improved Dataset for Entity Tracking in Texts
Zhang, Li
Xu, Hainiu
Kommula, Abhinav
Callison-Burch, Chris
Tandon, Niket
Computation and Language
Much text describes a changing world (e.g., procedures, stories, newswires), and understanding them requires tracking how entities change. An earlier dataset, OpenPI, provided crowdsourced annotations of entity state changes in text. However, a major limitation was that those annotations were free-form and did not identify salient changes, hampering model evaluation. To overcome these limitations, we present an improved dataset, OpenPI2.0, where entities and attributes are fully canonicalized and additional entity salience annotations are added. On our fairer evaluation setting, we find that current state-of-the-art language models are far from competent. We also show that using state changes of salient entities as a chain-of-thought prompt, downstream performance is improved on tasks such as question answering and classical planning, outperforming the setting involving all related entities indiscriminately. We offer OpenPI2.0 for the continued development of models that can understand the dynamics of entities in text.
title OpenPI2.0: An Improved Dataset for Entity Tracking in Texts
topic Computation and Language
url https://arxiv.org/abs/2305.14603