Sui Generis: Large Language Models for Authorship Attribution and Verification in Latin

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schmidt, Gleb, Gorovaia, Svetlana, Yamshchikov, Ivan P.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912069897420800
author Schmidt, Gleb
Gorovaia, Svetlana
Yamshchikov, Ivan P.
author_facet Schmidt, Gleb
Gorovaia, Svetlana
Yamshchikov, Ivan P.
contents This paper evaluates the performance of Large Language Models (LLMs) in authorship attribution and authorship verification tasks for Latin texts of the Patristic Era. The study showcases that LLMs can be robust in zero-shot authorship verification even on short texts without sophisticated feature engineering. Yet, the models can also be easily "mislead" by semantics. The experiments also demonstrate that steering the model's authorship analysis and decision-making is challenging, unlike what is reported in the studies dealing with high-resource modern languages. Although LLMs prove to be able to beat, under certain circumstances, the traditional baselines, obtaining a nuanced and truly explainable decision requires at best a lot of experimentation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_09245
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sui Generis: Large Language Models for Authorship Attribution and Verification in Latin
Schmidt, Gleb
Gorovaia, Svetlana
Yamshchikov, Ivan P.
Computation and Language
This paper evaluates the performance of Large Language Models (LLMs) in authorship attribution and authorship verification tasks for Latin texts of the Patristic Era. The study showcases that LLMs can be robust in zero-shot authorship verification even on short texts without sophisticated feature engineering. Yet, the models can also be easily "mislead" by semantics. The experiments also demonstrate that steering the model's authorship analysis and decision-making is challenging, unlike what is reported in the studies dealing with high-resource modern languages. Although LLMs prove to be able to beat, under certain circumstances, the traditional baselines, obtaining a nuanced and truly explainable decision requires at best a lot of experimentation.
title Sui Generis: Large Language Models for Authorship Attribution and Verification in Latin
topic Computation and Language
url https://arxiv.org/abs/2410.09245