CHATTER: A Character Attribution Dataset for Narrative Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baruah, Sabyasachee, Narayanan, Shrikanth
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916697457295360
author Baruah, Sabyasachee
Narayanan, Shrikanth
author_facet Baruah, Sabyasachee
Narayanan, Shrikanth
contents Computational narrative understanding studies the identification, description, and interaction of the elements of a narrative: characters, attributes, events, and relations. Narrative research has given considerable attention to defining and classifying character types. However, these character-type taxonomies do not generalize well because they are small, too simple, or specific to a domain. We require robust and reliable benchmarks to test whether narrative models truly understand the nuances of the character's development in the story. Our work addresses this by curating the CHATTER dataset that labels whether a character portrays some attribute for 88124 character-attribute pairs, encompassing 2998 characters, 12967 attributes and 660 movies. We validate a subset of CHATTER, called CHATTEREVAL, using human annotations to serve as a benchmark to evaluate the character attribution task in movie scripts. \evaldataset{} also assesses narrative understanding and the long-context modeling capacity of language models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05227
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CHATTER: A Character Attribution Dataset for Narrative Understanding
Baruah, Sabyasachee
Narayanan, Shrikanth
Computation and Language
Computational narrative understanding studies the identification, description, and interaction of the elements of a narrative: characters, attributes, events, and relations. Narrative research has given considerable attention to defining and classifying character types. However, these character-type taxonomies do not generalize well because they are small, too simple, or specific to a domain. We require robust and reliable benchmarks to test whether narrative models truly understand the nuances of the character's development in the story. Our work addresses this by curating the CHATTER dataset that labels whether a character portrays some attribute for 88124 character-attribute pairs, encompassing 2998 characters, 12967 attributes and 660 movies. We validate a subset of CHATTER, called CHATTEREVAL, using human annotations to serve as a benchmark to evaluate the character attribution task in movie scripts. \evaldataset{} also assesses narrative understanding and the long-context modeling capacity of language models.
title CHATTER: A Character Attribution Dataset for Narrative Understanding
topic Computation and Language
url https://arxiv.org/abs/2411.05227