M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Yuxia, Mansurov, Jonibek, Ivanov, Petar, Su, Jinyan, Shelmanov, Artem, Tsvigun, Akim, Whitehouse, Chenxi, Afzal, Osama Mohammed, Mahmoud, Tarek, Sasaki, Toru, Arnold, Thomas, Aji, Alham Fikri, Habash, Nizar, Gurevych, Iryna, Nakov, Preslav
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909132936708096
author Wang, Yuxia
Mansurov, Jonibek
Ivanov, Petar
Su, Jinyan
Shelmanov, Artem
Tsvigun, Akim
Whitehouse, Chenxi
Afzal, Osama Mohammed
Mahmoud, Tarek
Sasaki, Toru
Arnold, Thomas
Aji, Alham Fikri
Habash, Nizar
Gurevych, Iryna
Nakov, Preslav
author_facet Wang, Yuxia
Mansurov, Jonibek
Ivanov, Petar
Su, Jinyan
Shelmanov, Artem
Tsvigun, Akim
Whitehouse, Chenxi
Afzal, Osama Mohammed
Mahmoud, Tarek
Sasaki, Toru
Arnold, Thomas
Aji, Alham Fikri
Habash, Nizar
Gurevych, Iryna
Nakov, Preslav
contents Large language models (LLMs) have demonstrated remarkable capability to generate fluent responses to a wide variety of user queries. However, this has also raised concerns about the potential misuse of such texts in journalism, education, and academia. In this study, we strive to create automated systems that can detect machine-generated texts and pinpoint potential misuse. We first introduce a large-scale benchmark \textbf{M4}, which is a multi-generator, multi-domain, and multi-lingual corpus for machine-generated text detection. Through an extensive empirical study of this dataset, we show that it is challenging for detectors to generalize well on instances from unseen domains or LLMs. In such cases, detectors tend to misclassify machine-generated text as human-written. These results show that the problem is far from solved and that there is a lot of room for improvement. We believe that our dataset will enable future research towards more robust approaches to this pressing societal problem. The dataset is available at https://github.com/mbzuai-nlp/M4.
format Preprint
id arxiv_https___arxiv_org_abs_2305_14902
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection
Wang, Yuxia
Mansurov, Jonibek
Ivanov, Petar
Su, Jinyan
Shelmanov, Artem
Tsvigun, Akim
Whitehouse, Chenxi
Afzal, Osama Mohammed
Mahmoud, Tarek
Sasaki, Toru
Arnold, Thomas
Aji, Alham Fikri
Habash, Nizar
Gurevych, Iryna
Nakov, Preslav
Computation and Language
Large language models (LLMs) have demonstrated remarkable capability to generate fluent responses to a wide variety of user queries. However, this has also raised concerns about the potential misuse of such texts in journalism, education, and academia. In this study, we strive to create automated systems that can detect machine-generated texts and pinpoint potential misuse. We first introduce a large-scale benchmark \textbf{M4}, which is a multi-generator, multi-domain, and multi-lingual corpus for machine-generated text detection. Through an extensive empirical study of this dataset, we show that it is challenging for detectors to generalize well on instances from unseen domains or LLMs. In such cases, detectors tend to misclassify machine-generated text as human-written. These results show that the problem is far from solved and that there is a lot of room for improvement. We believe that our dataset will enable future research towards more robust approaches to this pressing societal problem. The dataset is available at https://github.com/mbzuai-nlp/M4.
title M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection
topic Computation and Language
url https://arxiv.org/abs/2305.14902