GenAI Content Detection Task 2: AI vs. Human -- Academic Essay Authenticity Challenge

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chowdhury, Shammur Absar, Almerekhi, Hind, Kutlu, Mucahid, Keles, Kaan Efe, Ahmad, Fatema, Mohiuddin, Tasnim, Mikros, George, Alam, Firoj
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929646302396416
author Chowdhury, Shammur Absar
Almerekhi, Hind
Kutlu, Mucahid
Keles, Kaan Efe
Ahmad, Fatema
Mohiuddin, Tasnim
Mikros, George
Alam, Firoj
author_facet Chowdhury, Shammur Absar
Almerekhi, Hind
Kutlu, Mucahid
Keles, Kaan Efe
Ahmad, Fatema
Mohiuddin, Tasnim
Mikros, George
Alam, Firoj
contents This paper presents a comprehensive overview of the first edition of the Academic Essay Authenticity Challenge, organized as part of the GenAI Content Detection shared tasks collocated with COLING 2025. This challenge focuses on detecting machine-generated vs. human-authored essays for academic purposes. The task is defined as follows: "Given an essay, identify whether it is generated by a machine or authored by a human.'' The challenge involves two languages: English and Arabic. During the evaluation phase, 25 teams submitted systems for English and 21 teams for Arabic, reflecting substantial interest in the task. Finally, seven teams submitted system description papers. The majority of submissions utilized fine-tuned transformer-based models, with one team employing Large Language Models (LLMs) such as Llama 2 and Llama 3. This paper outlines the task formulation, details the dataset construction process, and explains the evaluation framework. Additionally, we present a summary of the approaches adopted by participating teams. Nearly all submitted systems outperformed the n-gram-based baseline, with the top-performing systems achieving F1 scores exceeding 0.98 for both languages, indicating significant progress in the detection of machine-generated text.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18274
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GenAI Content Detection Task 2: AI vs. Human -- Academic Essay Authenticity Challenge
Chowdhury, Shammur Absar
Almerekhi, Hind
Kutlu, Mucahid
Keles, Kaan Efe
Ahmad, Fatema
Mohiuddin, Tasnim
Mikros, George
Alam, Firoj
Computation and Language
Artificial Intelligence
68T50
F.2.2; I.2.7
This paper presents a comprehensive overview of the first edition of the Academic Essay Authenticity Challenge, organized as part of the GenAI Content Detection shared tasks collocated with COLING 2025. This challenge focuses on detecting machine-generated vs. human-authored essays for academic purposes. The task is defined as follows: "Given an essay, identify whether it is generated by a machine or authored by a human.'' The challenge involves two languages: English and Arabic. During the evaluation phase, 25 teams submitted systems for English and 21 teams for Arabic, reflecting substantial interest in the task. Finally, seven teams submitted system description papers. The majority of submissions utilized fine-tuned transformer-based models, with one team employing Large Language Models (LLMs) such as Llama 2 and Llama 3. This paper outlines the task formulation, details the dataset construction process, and explains the evaluation framework. Additionally, we present a summary of the approaches adopted by participating teams. Nearly all submitted systems outperformed the n-gram-based baseline, with the top-performing systems achieving F1 scores exceeding 0.98 for both languages, indicating significant progress in the detection of machine-generated text.
title GenAI Content Detection Task 2: AI vs. Human -- Academic Essay Authenticity Challenge
topic Computation and Language
Artificial Intelligence
68T50
F.2.2; I.2.7
url https://arxiv.org/abs/2412.18274