GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dugan, Liam, Zhu, Andrew, Alam, Firoj, Nakov, Preslav, Apidianaki, Marianna, Callison-Burch, Chris
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912189830397952
author Dugan, Liam
Zhu, Andrew
Alam, Firoj
Nakov, Preslav
Apidianaki, Marianna
Callison-Burch, Chris
author_facet Dugan, Liam
Zhu, Andrew
Alam, Firoj
Nakov, Preslav
Apidianaki, Marianna
Callison-Burch, Chris
contents Recently there have been many shared tasks targeting the detection of generated text from Large Language Models (LLMs). However, these shared tasks tend to focus either on cases where text is limited to one particular domain or cases where text can be from many domains, some of which may not be seen during test time. In this shared task, using the newly released RAID benchmark, we aim to answer whether or not models can detect generated text from a large, yet fixed, number of domains and LLMs, all of which are seen during training. Over the course of three months, our task was attempted by 9 teams with 23 detector submissions. We find that multiple participants were able to obtain accuracies of over 99% on machine-generated text from RAID while maintaining a 5% False Positive Rate -- suggesting that detectors are able to robustly detect text from many domains and models simultaneously. We discuss potential interpretations of this result and provide directions for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2501_08913
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge
Dugan, Liam
Zhu, Andrew
Alam, Firoj
Nakov, Preslav
Apidianaki, Marianna
Callison-Burch, Chris
Computation and Language
Machine Learning
I.2.7
Recently there have been many shared tasks targeting the detection of generated text from Large Language Models (LLMs). However, these shared tasks tend to focus either on cases where text is limited to one particular domain or cases where text can be from many domains, some of which may not be seen during test time. In this shared task, using the newly released RAID benchmark, we aim to answer whether or not models can detect generated text from a large, yet fixed, number of domains and LLMs, all of which are seen during training. Over the course of three months, our task was attempted by 9 teams with 23 detector submissions. We find that multiple participants were able to obtain accuracies of over 99% on machine-generated text from RAID while maintaining a 5% False Positive Rate -- suggesting that detectors are able to robustly detect text from many domains and models simultaneously. We discuss potential interpretations of this result and provide directions for future research.
title GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge
topic Computation and Language
Machine Learning
I.2.7
url https://arxiv.org/abs/2501.08913