Saved in:
Bibliographic Details
Main Authors: Moon, Todd K, Gunther, Jacob H.
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2403.13253
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913273451905024
author Moon, Todd K
Gunther, Jacob H.
author_facet Moon, Todd K
Gunther, Jacob H.
contents Over the years there has been ongoing interest in detecting authorship of a text based on statistical properties of the text, such as by using occurrence rates of noncontextual words. In previous work, these techniques have been used, for example, to determine authorship of all of \emph{The Federalist Papers}. Such methods may be useful in more modern times to detect fake or AI authorship. Progress in statistical natural language parsers introduces the possibility of using grammatical structure to detect authorship. In this paper we explore a new possibility for detecting authorship using grammatical structural information extracted using a statistical natural language parser. This paper provides a proof of concept, testing author classification based on grammatical structure on a set of "proof texts," The Federalist Papers and Sanditon which have been as test cases in previous authorship detection studies. Several features extracted from the statistical natural language parser were explored: all subtrees of some depth from any level; rooted subtrees of some depth, part of speech, and part of speech by level in the parse tree. It was found to be helpful to project the features into a lower dimensional space. Statistical experiments on these documents demonstrate that information from a statistical parser can, in fact, assist in distinguishing authors.
format Preprint
id arxiv_https___arxiv_org_abs_2403_13253
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Document Author Classification Using Parsed Language Structure
Moon, Todd K
Gunther, Jacob H.
Computation and Language
Audio and Speech Processing
Over the years there has been ongoing interest in detecting authorship of a text based on statistical properties of the text, such as by using occurrence rates of noncontextual words. In previous work, these techniques have been used, for example, to determine authorship of all of \emph{The Federalist Papers}. Such methods may be useful in more modern times to detect fake or AI authorship. Progress in statistical natural language parsers introduces the possibility of using grammatical structure to detect authorship. In this paper we explore a new possibility for detecting authorship using grammatical structural information extracted using a statistical natural language parser. This paper provides a proof of concept, testing author classification based on grammatical structure on a set of "proof texts," The Federalist Papers and Sanditon which have been as test cases in previous authorship detection studies. Several features extracted from the statistical natural language parser were explored: all subtrees of some depth from any level; rooted subtrees of some depth, part of speech, and part of speech by level in the parse tree. It was found to be helpful to project the features into a lower dimensional space. Statistical experiments on these documents demonstrate that information from a statistical parser can, in fact, assist in distinguishing authors.
title Document Author Classification Using Parsed Language Structure
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2403.13253