The Elephant in the Coreference Room: Resolving Coreference in Full-Length French Fiction Works

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bourgois, Antoine, Poibeau, Thierry
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915559495434240
author Bourgois, Antoine
Poibeau, Thierry
author_facet Bourgois, Antoine
Poibeau, Thierry
contents While coreference resolution is attracting more interest than ever from computational literature researchers, representative datasets of fully annotated long documents remain surprisingly scarce. In this paper, we introduce a new annotated corpus of three full-length French novels, totaling over 285,000 tokens. Unlike previous datasets focused on shorter texts, our corpus addresses the challenges posed by long, complex literary works, enabling evaluation of coreference models in the context of long reference chains. We present a modular coreference resolution pipeline that allows for fine-grained error analysis. We show that our approach is competitive and scales effectively to long documents. Finally, we demonstrate its usefulness to infer the gender of fictional characters, showcasing its relevance for both literary analysis and downstream NLP tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15594
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Elephant in the Coreference Room: Resolving Coreference in Full-Length French Fiction Works
Bourgois, Antoine
Poibeau, Thierry
Computation and Language
While coreference resolution is attracting more interest than ever from computational literature researchers, representative datasets of fully annotated long documents remain surprisingly scarce. In this paper, we introduce a new annotated corpus of three full-length French novels, totaling over 285,000 tokens. Unlike previous datasets focused on shorter texts, our corpus addresses the challenges posed by long, complex literary works, enabling evaluation of coreference models in the context of long reference chains. We present a modular coreference resolution pipeline that allows for fine-grained error analysis. We show that our approach is competitive and scales effectively to long documents. Finally, we demonstrate its usefulness to infer the gender of fictional characters, showcasing its relevance for both literary analysis and downstream NLP tasks.
title The Elephant in the Coreference Room: Resolving Coreference in Full-Length French Fiction Works
topic Computation and Language
url https://arxiv.org/abs/2510.15594