Separating Style from Substance: Enhancing Cross-Genre Authorship Attribution through Data Selection and Presentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fincke, Steven, Boschee, Elizabeth
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913463282958336
author Fincke, Steven
Boschee, Elizabeth
author_facet Fincke, Steven
Boschee, Elizabeth
contents The task of deciding whether two documents are written by the same author is challenging for both machines and humans. This task is even more challenging when the two documents are written about different topics (e.g. baseball vs. politics) or in different genres (e.g. a blog post vs. an academic article). For machines, the problem is complicated by the relative lack of real-world training examples that cross the topic boundary and the vanishing scarcity of cross-genre data. We propose targeted methods for training data selection and a novel learning curriculum that are designed to discourage a model's reliance on topic information for authorship attribution and correspondingly force it to incorporate information more robustly indicative of style no matter the topic. These refinements yield a 62.7% relative improvement in average cross-genre authorship attribution, as well as 16.6% in the per-genre condition.
format Preprint
id arxiv_https___arxiv_org_abs_2408_05192
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Separating Style from Substance: Enhancing Cross-Genre Authorship Attribution through Data Selection and Presentation
Fincke, Steven
Boschee, Elizabeth
Computation and Language
The task of deciding whether two documents are written by the same author is challenging for both machines and humans. This task is even more challenging when the two documents are written about different topics (e.g. baseball vs. politics) or in different genres (e.g. a blog post vs. an academic article). For machines, the problem is complicated by the relative lack of real-world training examples that cross the topic boundary and the vanishing scarcity of cross-genre data. We propose targeted methods for training data selection and a novel learning curriculum that are designed to discourage a model's reliance on topic information for authorship attribution and correspondingly force it to incorporate information more robustly indicative of style no matter the topic. These refinements yield a 62.7% relative improvement in average cross-genre authorship attribution, as well as 16.6% in the per-genre condition.
title Separating Style from Substance: Enhancing Cross-Genre Authorship Attribution through Data Selection and Presentation
topic Computation and Language
url https://arxiv.org/abs/2408.05192