CEAID: Benchmark of Multilingual Machine-Generated Text Detection Methods for Central European Languages

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Macko, Dominik, Kopal, Jakub
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915524815880192
author Macko, Dominik
Kopal, Jakub
author_facet Macko, Dominik
Kopal, Jakub
contents Machine-generated text detection, as an important task, is predominantly focused on English in research. This makes the existing detectors almost unusable for non-English languages, relying purely on cross-lingual transferability. There exist only a few works focused on any of Central European languages, leaving the transferability towards these languages rather unexplored. We fill this gap by providing the first benchmark of detection methods focused on this region, while also providing comparison of train-languages combinations to identify the best performing ones. We focus on multi-domain, multi-generator, and multilingual evaluation, pinpointing the differences of individual aspects, as well as adversarial robustness of detection methods. Supervised finetuned detectors in the Central European languages are found the most performant in these languages as well as the most resistant against obfuscation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_26051
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CEAID: Benchmark of Multilingual Machine-Generated Text Detection Methods for Central European Languages
Macko, Dominik
Kopal, Jakub
Computation and Language
Artificial Intelligence
Machine-generated text detection, as an important task, is predominantly focused on English in research. This makes the existing detectors almost unusable for non-English languages, relying purely on cross-lingual transferability. There exist only a few works focused on any of Central European languages, leaving the transferability towards these languages rather unexplored. We fill this gap by providing the first benchmark of detection methods focused on this region, while also providing comparison of train-languages combinations to identify the best performing ones. We focus on multi-domain, multi-generator, and multilingual evaluation, pinpointing the differences of individual aspects, as well as adversarial robustness of detection methods. Supervised finetuned detectors in the Central European languages are found the most performant in these languages as well as the most resistant against obfuscation.
title CEAID: Benchmark of Multilingual Machine-Generated Text Detection Methods for Central European Languages
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.26051