Cross-Genre Authorship Attribution via LLM-Based Retrieve-and-Rerank

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Agarwal, Shantanu, Barry, Joel, Fincke, Steven, Miller, Scott
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911220171276288
author Agarwal, Shantanu
Barry, Joel
Fincke, Steven
Miller, Scott
author_facet Agarwal, Shantanu
Barry, Joel
Fincke, Steven
Miller, Scott
contents Authorship attribution (AA) is the task of identifying the most likely author of a query document from a predefined set of candidate authors. We introduce a two-stage retrieve-and-rerank framework that finetunes LLMs for cross-genre AA. Unlike the field of information retrieval (IR), where retrieve-and-rerank is a de facto strategy, cross-genre AA systems must avoid relying on topical cues and instead learn to identify author-specific linguistic patterns that are independent of the text's subject matter (genre/domain/topic). Consequently, for the reranker, we demonstrate that training strategies commonly used in IR are fundamentally misaligned with cross-genre AA, leading to suboptimal behavior. To address this, we introduce a targeted data curation strategy that enables the reranker to effectively learn author-discriminative signals. Using our LLM-based retrieve-and-rerank pipeline, we achieve substantial gains of 22.3 and 34.4 absolute Success@8 points over the previous state-of-the-art on HIATUS's challenging HRS1 and HRS2 cross-genre AA benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16819
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cross-Genre Authorship Attribution via LLM-Based Retrieve-and-Rerank
Agarwal, Shantanu
Barry, Joel
Fincke, Steven
Miller, Scott
Computation and Language
Authorship attribution (AA) is the task of identifying the most likely author of a query document from a predefined set of candidate authors. We introduce a two-stage retrieve-and-rerank framework that finetunes LLMs for cross-genre AA. Unlike the field of information retrieval (IR), where retrieve-and-rerank is a de facto strategy, cross-genre AA systems must avoid relying on topical cues and instead learn to identify author-specific linguistic patterns that are independent of the text's subject matter (genre/domain/topic). Consequently, for the reranker, we demonstrate that training strategies commonly used in IR are fundamentally misaligned with cross-genre AA, leading to suboptimal behavior. To address this, we introduce a targeted data curation strategy that enables the reranker to effectively learn author-discriminative signals. Using our LLM-based retrieve-and-rerank pipeline, we achieve substantial gains of 22.3 and 34.4 absolute Success@8 points over the previous state-of-the-art on HIATUS's challenging HRS1 and HRS2 cross-genre AA benchmarks.
title Cross-Genre Authorship Attribution via LLM-Based Retrieve-and-Rerank
topic Computation and Language
url https://arxiv.org/abs/2510.16819