Contrasting Linguistic Patterns in Human and LLM-Generated News Text

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Muñoz-Ortiz, Alberto, Gómez-Rodríguez, Carlos, Vilares, David
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917764434755584
author Muñoz-Ortiz, Alberto
Gómez-Rodríguez, Carlos
Vilares, David
author_facet Muñoz-Ortiz, Alberto
Gómez-Rodríguez, Carlos
Vilares, David
contents We conduct a quantitative analysis contrasting human-written English news text with comparable large language model (LLM) output from six different LLMs that cover three different families and four sizes in total. Our analysis spans several measurable linguistic dimensions, including morphological, syntactic, psychometric, and sociolinguistic aspects. The results reveal various measurable differences between human and AI-generated texts. Human texts exhibit more scattered sentence length distributions, more variety of vocabulary, a distinct use of dependency and constituent types, shorter constituents, and more optimized dependency distances. Humans tend to exhibit stronger negative emotions (such as fear and disgust) and less joy compared to text generated by LLMs, with the toxicity of these models increasing as their size grows. LLM outputs use more numbers, symbols and auxiliaries (suggesting objective language) than human texts, as well as more pronouns. The sexist bias prevalent in human text is also expressed by LLMs, and even magnified in all of them but one. Differences between LLMs and humans are larger than between LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2308_09067
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Contrasting Linguistic Patterns in Human and LLM-Generated News Text
Muñoz-Ortiz, Alberto
Gómez-Rodríguez, Carlos
Vilares, David
Computation and Language
68T50
I.2.7
We conduct a quantitative analysis contrasting human-written English news text with comparable large language model (LLM) output from six different LLMs that cover three different families and four sizes in total. Our analysis spans several measurable linguistic dimensions, including morphological, syntactic, psychometric, and sociolinguistic aspects. The results reveal various measurable differences between human and AI-generated texts. Human texts exhibit more scattered sentence length distributions, more variety of vocabulary, a distinct use of dependency and constituent types, shorter constituents, and more optimized dependency distances. Humans tend to exhibit stronger negative emotions (such as fear and disgust) and less joy compared to text generated by LLMs, with the toxicity of these models increasing as their size grows. LLM outputs use more numbers, symbols and auxiliaries (suggesting objective language) than human texts, as well as more pronouns. The sexist bias prevalent in human text is also expressed by LLMs, and even magnified in all of them but one. Differences between LLMs and humans are larger than between LLMs.
title Contrasting Linguistic Patterns in Human and LLM-Generated News Text
topic Computation and Language
68T50
I.2.7
url https://arxiv.org/abs/2308.09067