On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Geng, Mingmeng, Poibeau, Thierry
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917037773684736
author Geng, Mingmeng
Poibeau, Thierry
author_facet Geng, Mingmeng
Poibeau, Thierry
contents With the widespread use of large language models (LLMs), many researchers have turned their attention to detecting text generated by them. However, there is no consistent or precise definition of their target, namely "LLM-generated text". Differences in usage scenarios and the diversity of LLMs further increase the difficulty of detection. What is commonly regarded as the detecting target usually represents only a subset of the text that LLMs can potentially produce. Human edits to LLM outputs, together with the subtle influences that LLMs exert on their users, are blurring the line between LLM-generated and human-written text. Existing benchmarks and evaluation approaches do not adequately address the various conditions in real-world detector applications. Hence, the numerical results of detectors are often misunderstood, and their significance is diminishing. Therefore, detectors remain useful under specific conditions, but their results should be interpreted only as references rather than decisive indicators.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20810
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?
Geng, Mingmeng
Poibeau, Thierry
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
With the widespread use of large language models (LLMs), many researchers have turned their attention to detecting text generated by them. However, there is no consistent or precise definition of their target, namely "LLM-generated text". Differences in usage scenarios and the diversity of LLMs further increase the difficulty of detection. What is commonly regarded as the detecting target usually represents only a subset of the text that LLMs can potentially produce. Human edits to LLM outputs, together with the subtle influences that LLMs exert on their users, are blurring the line between LLM-generated and human-written text. Existing benchmarks and evaluation approaches do not adequately address the various conditions in real-world detector applications. Hence, the numerical results of detectors are often misunderstood, and their significance is diminishing. Therefore, detectors remain useful under specific conditions, but their results should be interpreted only as references rather than decisive indicators.
title On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2510.20810