Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Schelb, Julian, Wittweiler, Michael, Revellio, Marie, Feichtinger, Barbara, Spitz, Andreas
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910004805632000
author Schelb, Julian
Wittweiler, Michael
Revellio, Marie
Feichtinger, Barbara
Spitz, Andreas
author_facet Schelb, Julian
Wittweiler, Michael
Revellio, Marie
Feichtinger, Barbara
Spitz, Andreas
contents Tracing connections between historical texts is an important part of intertextual research, enabling scholars to reconstruct the virtual library of a writer and identify the sources influencing their creative process. These intertextual links manifest in diverse forms, ranging from direct verbatim quotations to subtle allusions and paraphrases disguised by morphological variation. Language models offer a promising path forward due to their capability of capturing semantic similarity beyond lexical overlap. However, the development of new methods for this task is held back by the scarcity of standardized benchmarks and easy-to-use datasets. We address this gap by introducing Loci Similes, a benchmark for Latin intertextuality detection comprising of a curated dataset of ~172k text segments containing 545 expert-verified parallels linking Late Antique authors to a corpus of classical authors. Using this data, we establish baselines for retrieval and classification of intertextualities with state-of-the-art LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2601_07533
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature
Schelb, Julian
Wittweiler, Michael
Revellio, Marie
Feichtinger, Barbara
Spitz, Andreas
Information Retrieval
Computation and Language
Digital Libraries
Tracing connections between historical texts is an important part of intertextual research, enabling scholars to reconstruct the virtual library of a writer and identify the sources influencing their creative process. These intertextual links manifest in diverse forms, ranging from direct verbatim quotations to subtle allusions and paraphrases disguised by morphological variation. Language models offer a promising path forward due to their capability of capturing semantic similarity beyond lexical overlap. However, the development of new methods for this task is held back by the scarcity of standardized benchmarks and easy-to-use datasets. We address this gap by introducing Loci Similes, a benchmark for Latin intertextuality detection comprising of a curated dataset of ~172k text segments containing 545 expert-verified parallels linking Late Antique authors to a corpus of classical authors. Using this data, we establish baselines for retrieval and classification of intertextualities with state-of-the-art LLMs.
title Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature
topic Information Retrieval
Computation and Language
Digital Libraries
url https://arxiv.org/abs/2601.07533