Benchmark of stylistic variation in LLM-generated texts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Milička, Jiří, Marklová, Anna, Cvrček, Václav
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916956974612480
author Milička, Jiří
Marklová, Anna
Cvrček, Václav
author_facet Milička, Jiří
Marklová, Anna
Cvrček, Václav
contents This study investigates the register variation in texts written by humans and comparable texts produced by large language models (LLMs). Biber's multidimensional analysis (MDA) is applied to a sample of human-written texts and AI-created texts generated to be their counterparts to find the dimensions of variation in which LLMs differ most significantly and most systematically from humans. As textual material, a new LLM-generated corpus AI-Brown is used, which is comparable to BE-21 (a Brown family corpus representing contemporary British English). Since all languages except English are underrepresented in the training data of frontier LLMs, similar analysis is replicated on Czech using AI-Koditex corpus and Czech multidimensional model. Examined were 16 frontier models in various settings and prompts, with emphasis placed on the difference between base models and instruction-tuned models. Based on this, a benchmark is created through which models can be compared with each other and ranked in interpretable dimensions.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10179
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmark of stylistic variation in LLM-generated texts
Milička, Jiří
Marklová, Anna
Cvrček, Václav
Computation and Language
Artificial Intelligence
This study investigates the register variation in texts written by humans and comparable texts produced by large language models (LLMs). Biber's multidimensional analysis (MDA) is applied to a sample of human-written texts and AI-created texts generated to be their counterparts to find the dimensions of variation in which LLMs differ most significantly and most systematically from humans. As textual material, a new LLM-generated corpus AI-Brown is used, which is comparable to BE-21 (a Brown family corpus representing contemporary British English). Since all languages except English are underrepresented in the training data of frontier LLMs, similar analysis is replicated on Czech using AI-Koditex corpus and Czech multidimensional model. Examined were 16 frontier models in various settings and prompts, with emphasis placed on the difference between base models and instruction-tuned models. Based on this, a benchmark is created through which models can be compared with each other and ranked in interpretable dimensions.
title Benchmark of stylistic variation in LLM-generated texts
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.10179