MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Macko, Dominik, Kopal, Jakub, Moro, Robert, Srba, Ivan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913959359021056
author Macko, Dominik
Kopal, Jakub
Moro, Robert
Srba, Ivan
author_facet Macko, Dominik
Kopal, Jakub
Moro, Robert
Srba, Ivan
contents Recent LLMs are able to generate high-quality multilingual texts, indistinguishable for humans from authentic human-written ones. Research in machine-generated text detection is however mostly focused on the English language and longer texts, such as news articles, scientific papers or student essays. Social-media texts are usually much shorter and often feature informal language, grammatical errors, or distinct linguistic items (e.g., emoticons, hashtags). There is a gap in studying the ability of existing methods in detection of such texts, reflected also in the lack of existing multilingual benchmark datasets. To fill this gap we propose the first multilingual (22 languages) and multi-platform (5 social media platforms) dataset for benchmarking machine-generated text detection in the social-media domain, called MultiSocial. It contains 472,097 texts, of which about 58k are human-written and approximately the same amount is generated by each of 7 multilingual LLMs. We use this benchmark to compare existing detection methods in zero-shot as well as fine-tuned form. Our results indicate that the fine-tuned detectors have no problem to be trained on social-media texts and that the platform selection for training matters.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12549
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts
Macko, Dominik
Kopal, Jakub
Moro, Robert
Srba, Ivan
Computation and Language
Artificial Intelligence
Recent LLMs are able to generate high-quality multilingual texts, indistinguishable for humans from authentic human-written ones. Research in machine-generated text detection is however mostly focused on the English language and longer texts, such as news articles, scientific papers or student essays. Social-media texts are usually much shorter and often feature informal language, grammatical errors, or distinct linguistic items (e.g., emoticons, hashtags). There is a gap in studying the ability of existing methods in detection of such texts, reflected also in the lack of existing multilingual benchmark datasets. To fill this gap we propose the first multilingual (22 languages) and multi-platform (5 social media platforms) dataset for benchmarking machine-generated text detection in the social-media domain, called MultiSocial. It contains 472,097 texts, of which about 58k are human-written and approximately the same amount is generated by each of 7 multilingual LLMs. We use this benchmark to compare existing detection methods in zero-shot as well as fine-tuned form. Our results indicate that the fine-tuned detectors have no problem to be trained on social-media texts and that the platform selection for training matters.
title MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2406.12549