Saved in:
Bibliographic Details
Main Authors: Bizzoni, Yuri, Feldkamp, Pascale, Lassen, Ida Marie, Jacobsen, Mia, Thomsen, Mads Rosendahl, Nielbo, Kristoffer
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2404.04022
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909168939565056
author Bizzoni, Yuri
Feldkamp, Pascale
Lassen, Ida Marie
Jacobsen, Mia
Thomsen, Mads Rosendahl
Nielbo, Kristoffer
author_facet Bizzoni, Yuri
Feldkamp, Pascale
Lassen, Ida Marie
Jacobsen, Mia
Thomsen, Mads Rosendahl
Nielbo, Kristoffer
contents In this study, we employ a classification approach to show that different categories of literary "quality" display unique linguistic profiles, leveraging a corpus that encompasses titles from the Norton Anthology, Penguin Classics series, and the Open Syllabus project, contrasted against contemporary bestsellers, Nobel prize winners and recipients of prestigious literary awards. Our analysis reveals that canonical and so called high-brow texts exhibit distinct textual features when compared to other quality categories such as bestsellers and popular titles as well as to control groups, likely responding to distinct (but not mutually exclusive) models of quality. We apply a classic machine learning approach, namely Random Forest, to distinguish quality novels from "control groups", achieving up to 77\% F1 scores in differentiating between the categories. We find that quality category tend to be easier to distinguish from control groups than from other quality categories, suggesting than literary quality features might be distinguishable but shared through quality proxies.
format Preprint
id arxiv_https___arxiv_org_abs_2404_04022
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Good Books are Complex Matters: Gauging Complexity Profiles Across Diverse Categories of Perceived Literary Quality
Bizzoni, Yuri
Feldkamp, Pascale
Lassen, Ida Marie
Jacobsen, Mia
Thomsen, Mads Rosendahl
Nielbo, Kristoffer
Computation and Language
In this study, we employ a classification approach to show that different categories of literary "quality" display unique linguistic profiles, leveraging a corpus that encompasses titles from the Norton Anthology, Penguin Classics series, and the Open Syllabus project, contrasted against contemporary bestsellers, Nobel prize winners and recipients of prestigious literary awards. Our analysis reveals that canonical and so called high-brow texts exhibit distinct textual features when compared to other quality categories such as bestsellers and popular titles as well as to control groups, likely responding to distinct (but not mutually exclusive) models of quality. We apply a classic machine learning approach, namely Random Forest, to distinguish quality novels from "control groups", achieving up to 77\% F1 scores in differentiating between the categories. We find that quality category tend to be easier to distinguish from control groups than from other quality categories, suggesting than literary quality features might be distinguishable but shared through quality proxies.
title Good Books are Complex Matters: Gauging Complexity Profiles Across Diverse Categories of Perceived Literary Quality
topic Computation and Language
url https://arxiv.org/abs/2404.04022