How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nieth, Björn, Gracheva, Marianna, Mahlberg, Michaela, Eskofier, Bjoern, Salin, Emmanuelle
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918522662158336
author Nieth, Björn
Gracheva, Marianna
Mahlberg, Michaela
Eskofier, Bjoern
Salin, Emmanuelle
author_facet Nieth, Björn
Gracheva, Marianna
Mahlberg, Michaela
Eskofier, Bjoern
Salin, Emmanuelle
contents While factual correctness and task-performance have been in focus of Large Language Model (LLM) research for a long time, the fundamental question of how human-like generated texts are on a linguistic level has been underexplored. From a corpus-linguistic perspective, language production is inherently context-dependent, with distinct communicative contexts giving rise to differences in frequencies and co-occurrence patterns of linguistic features. A text failing to adhere to these patterns can be content-wise correct, but still be unfavorable to human readers. In this work, we propose a context-aware evaluation framework in which human-likeness is assessed using a two-sample problem between the linguistic feature distribution of a human reference corpus for a given register and a corresponding LLM-generated corpus. We implement this framework using the Maximum Mean Discrepancy (MMD) and the 67 lexico-grammatical features introduced by Biber, which are commonly applied in corpus linguistics. In our experiments, we compare seven instruction-tuned, open-source models across five English-language datasets spanning distinct registers against a human baseline. While across all tested setups, LLMs deviate from the human baseline, which models are closest to human language depends on the register and is not dictated by model size.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23651
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
Nieth, Björn
Gracheva, Marianna
Mahlberg, Michaela
Eskofier, Bjoern
Salin, Emmanuelle
Computation and Language
68T50,
I.2.7; I.2.6
While factual correctness and task-performance have been in focus of Large Language Model (LLM) research for a long time, the fundamental question of how human-like generated texts are on a linguistic level has been underexplored. From a corpus-linguistic perspective, language production is inherently context-dependent, with distinct communicative contexts giving rise to differences in frequencies and co-occurrence patterns of linguistic features. A text failing to adhere to these patterns can be content-wise correct, but still be unfavorable to human readers. In this work, we propose a context-aware evaluation framework in which human-likeness is assessed using a two-sample problem between the linguistic feature distribution of a human reference corpus for a given register and a corresponding LLM-generated corpus. We implement this framework using the Maximum Mean Discrepancy (MMD) and the 67 lexico-grammatical features introduced by Biber, which are commonly applied in corpus linguistics. In our experiments, we compare seven instruction-tuned, open-source models across five English-language datasets spanning distinct registers against a human baseline. While across all tested setups, LLMs deviate from the human baseline, which models are closest to human language depends on the register and is not dictated by model size.
title How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
topic Computation and Language
68T50,
I.2.7; I.2.6
url https://arxiv.org/abs/2605.23651