Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sarkar, Sujoy, Sarkar, Gourav, Jagadeeshan, Manoj Balaji, Sandhan, Jivnesh, Krishna, Amrith, Goyal, Pawan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912602493288448
author Sarkar, Sujoy
Sarkar, Gourav
Jagadeeshan, Manoj Balaji
Sandhan, Jivnesh
Krishna, Amrith
Goyal, Pawan
author_facet Sarkar, Sujoy
Sarkar, Gourav
Jagadeeshan, Manoj Balaji
Sandhan, Jivnesh
Krishna, Amrith
Goyal, Pawan
contents High lexical variation, ambiguous references, and long-range dependencies make entity resolution in literary texts particularly challenging. We present Mahānāma, the first large-scale dataset for end-to-end Entity Discovery and Linking (EDL) in Sanskrit, a morphologically rich and under-resourced language. Derived from the Mahābhārata, the world's longest epic, the dataset comprises over 109K named entity mentions mapped to 5.5K unique entities, and is aligned with an English knowledge base to support cross-lingual linking. The complex narrative structure of Mahānāma, coupled with extensive name variation and ambiguity, poses significant challenges to resolution systems. Our evaluation reveals that current coreference and entity linking models struggle when evaluated on the global context of the test set. These results highlight the limitations of current approaches in resolving entities within such complex discourse. Mahānāma thus provides a unique benchmark for advancing entity resolution, especially in literary domains.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19844
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking
Sarkar, Sujoy
Sarkar, Gourav
Jagadeeshan, Manoj Balaji
Sandhan, Jivnesh
Krishna, Amrith
Goyal, Pawan
Computation and Language
High lexical variation, ambiguous references, and long-range dependencies make entity resolution in literary texts particularly challenging. We present Mahānāma, the first large-scale dataset for end-to-end Entity Discovery and Linking (EDL) in Sanskrit, a morphologically rich and under-resourced language. Derived from the Mahābhārata, the world's longest epic, the dataset comprises over 109K named entity mentions mapped to 5.5K unique entities, and is aligned with an English knowledge base to support cross-lingual linking. The complex narrative structure of Mahānāma, coupled with extensive name variation and ambiguity, poses significant challenges to resolution systems. Our evaluation reveals that current coreference and entity linking models struggle when evaluated on the global context of the test set. These results highlight the limitations of current approaches in resolving entities within such complex discourse. Mahānāma thus provides a unique benchmark for advancing entity resolution, especially in literary domains.
title Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking
topic Computation and Language
url https://arxiv.org/abs/2509.19844