MAGE: Machine-generated Text Detection in the Wild

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yafu, Li, Qintong, Cui, Leyang, Bi, Wei, Wang, Zhilin, Wang, Longyue, Yang, Linyi, Shi, Shuming, Zhang, Yue
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929349455773696
author Li, Yafu
Li, Qintong
Cui, Leyang
Bi, Wei
Wang, Zhilin
Wang, Longyue
Yang, Linyi
Shi, Shuming
Zhang, Yue
author_facet Li, Yafu
Li, Qintong
Cui, Leyang
Bi, Wei
Wang, Zhilin
Wang, Longyue
Yang, Linyi
Shi, Shuming
Zhang, Yue
contents Large language models (LLMs) have achieved human-level text generation, emphasizing the need for effective AI-generated text detection to mitigate risks like the spread of fake news and plagiarism. Existing research has been constrained by evaluating detection methods on specific domains or particular language models. In practical scenarios, however, the detector faces texts from various domains or LLMs without knowing their sources. To this end, we build a comprehensive testbed by gathering texts from diverse human writings and texts generated by different LLMs. Empirical results show challenges in distinguishing machine-generated texts from human-authored ones across various scenarios, especially out-of-distribution. These challenges are due to the decreasing linguistic distinctions between the two sources. Despite challenges, the top-performing detector can identify 86.54% out-of-domain texts generated by a new LLM, indicating the feasibility for application scenarios. We release our resources at https://github.com/yafuly/MAGE.
format Preprint
id arxiv_https___arxiv_org_abs_2305_13242
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MAGE: Machine-generated Text Detection in the Wild
Li, Yafu
Li, Qintong
Cui, Leyang
Bi, Wei
Wang, Zhilin
Wang, Longyue
Yang, Linyi
Shi, Shuming
Zhang, Yue
Computation and Language
Large language models (LLMs) have achieved human-level text generation, emphasizing the need for effective AI-generated text detection to mitigate risks like the spread of fake news and plagiarism. Existing research has been constrained by evaluating detection methods on specific domains or particular language models. In practical scenarios, however, the detector faces texts from various domains or LLMs without knowing their sources. To this end, we build a comprehensive testbed by gathering texts from diverse human writings and texts generated by different LLMs. Empirical results show challenges in distinguishing machine-generated texts from human-authored ones across various scenarios, especially out-of-distribution. These challenges are due to the decreasing linguistic distinctions between the two sources. Despite challenges, the top-performing detector can identify 86.54% out-of-domain texts generated by a new LLM, indicating the feasibility for application scenarios. We release our resources at https://github.com/yafuly/MAGE.
title MAGE: Machine-generated Text Detection in the Wild
topic Computation and Language
url https://arxiv.org/abs/2305.13242