Atomic Literary Styling: Mechanistic Manipulation of Prose Generation in Neural Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autore principale: Enkhbayar, Tsogt-Ochir
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911222593486848
author Enkhbayar, Tsogt-Ochir
author_facet Enkhbayar, Tsogt-Ochir
contents We present a mechanistic analysis of literary style in GPT-2, identifying individual neurons that discriminate between exemplary prose and rigid AI-generated text. Using Herman Melville's Bartleby, the Scrivener as a corpus, we extract activation patterns from 355 million parameters across 32,768 neurons in late layers. We find 27,122 statistically significant discriminative neurons ($p < 0.05$), with effect sizes up to $|d| = 1.4$. Through systematic ablation studies, we discover a paradoxical result: while these neurons correlate with literary text during analysis, removing them often improves rather than degrades generated prose quality. Specifically, ablating 50 high-discriminating neurons yields a 25.7% improvement in literary style metrics. This demonstrates a critical gap between observational correlation and causal necessity in neural networks. Our findings challenge the assumption that neurons which activate on desirable inputs will produce those outputs during generation, with implications for mechanistic interpretability research and AI alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17909
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Atomic Literary Styling: Mechanistic Manipulation of Prose Generation in Neural Language Models
Enkhbayar, Tsogt-Ochir
Computation and Language
We present a mechanistic analysis of literary style in GPT-2, identifying individual neurons that discriminate between exemplary prose and rigid AI-generated text. Using Herman Melville's Bartleby, the Scrivener as a corpus, we extract activation patterns from 355 million parameters across 32,768 neurons in late layers. We find 27,122 statistically significant discriminative neurons ($p < 0.05$), with effect sizes up to $|d| = 1.4$. Through systematic ablation studies, we discover a paradoxical result: while these neurons correlate with literary text during analysis, removing them often improves rather than degrades generated prose quality. Specifically, ablating 50 high-discriminating neurons yields a 25.7% improvement in literary style metrics. This demonstrates a critical gap between observational correlation and causal necessity in neural networks. Our findings challenge the assumption that neurons which activate on desirable inputs will produce those outputs during generation, with implications for mechanistic interpretability research and AI alignment.
title Atomic Literary Styling: Mechanistic Manipulation of Prose Generation in Neural Language Models
topic Computation and Language
url https://arxiv.org/abs/2510.17909