A Comprehensive Dataset for Human vs. AI Generated Text Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Roy, Rajarshi, Singh, Gurpreet, Aziz, Ashhar, Bajpai, Shashwat, Imanpour, Nasrin, Biswas, Shwetangshu, Wanaskar, Kapil, Patwa, Parth, Ghosh, Subhankar, Dixit, Shreyas, Pal, Nilesh Ranjan, Rawte, Vipula, Garimella, Ritvik, Jena, Gaytri, Das, Amitava, Sheth, Amit, Sharma, Vasu, Reganti, Aishwarya Naresh, Jain, Vinija, Chadha, Aman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914598503841792
author Roy, Rajarshi
Singh, Gurpreet
Aziz, Ashhar
Bajpai, Shashwat
Imanpour, Nasrin
Biswas, Shwetangshu
Wanaskar, Kapil
Patwa, Parth
Ghosh, Subhankar
Dixit, Shreyas
Pal, Nilesh Ranjan
Rawte, Vipula
Garimella, Ritvik
Jena, Gaytri
Das, Amitava
Sheth, Amit
Sharma, Vasu
Reganti, Aishwarya Naresh
Jain, Vinija
Chadha, Aman
author_facet Roy, Rajarshi
Singh, Gurpreet
Aziz, Ashhar
Bajpai, Shashwat
Imanpour, Nasrin
Biswas, Shwetangshu
Wanaskar, Kapil
Patwa, Parth
Ghosh, Subhankar
Dixit, Shreyas
Pal, Nilesh Ranjan
Rawte, Vipula
Garimella, Ritvik
Jena, Gaytri
Das, Amitava
Sheth, Amit
Sharma, Vasu
Reganti, Aishwarya Naresh
Jain, Vinija
Chadha, Aman
contents The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authenticity, misinformation, and trustworthiness. Addressing the challenge of reliably detecting AI-generated text and attributing it to specific models requires large-scale, diverse, and well-annotated datasets. In this work, we present a comprehensive dataset comprising over 73,193 text samples that combine authentic New York Times articles with synthetic versions generated by multiple state-of-the-art LLMs including Gemma-2-9b, Mistral-7B, Qwen-2-72B, LLaMA-8B, Yi-Large, and GPT-4-o. The dataset provides original article abstracts as prompts, full human-authored narratives. We establish baseline results for two key tasks: distinguishing human-written from AI-generated text, achieving an accuracy of 58.35\%, and attributing AI texts to their generating models with an accuracy of 8.92\%. By bridging real-world journalistic content with modern generative models, the dataset aims to catalyze the development of robust detection and attribution methods, fostering trust and transparency in the era of generative AI. Our dataset is available at: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Text_Dataset
format Preprint
id arxiv_https___arxiv_org_abs_2510_22874
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Comprehensive Dataset for Human vs. AI Generated Text Detection
Roy, Rajarshi
Singh, Gurpreet
Aziz, Ashhar
Bajpai, Shashwat
Imanpour, Nasrin
Biswas, Shwetangshu
Wanaskar, Kapil
Patwa, Parth
Ghosh, Subhankar
Dixit, Shreyas
Pal, Nilesh Ranjan
Rawte, Vipula
Garimella, Ritvik
Jena, Gaytri
Das, Amitava
Sheth, Amit
Sharma, Vasu
Reganti, Aishwarya Naresh
Jain, Vinija
Chadha, Aman
Computation and Language
The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authenticity, misinformation, and trustworthiness. Addressing the challenge of reliably detecting AI-generated text and attributing it to specific models requires large-scale, diverse, and well-annotated datasets. In this work, we present a comprehensive dataset comprising over 73,193 text samples that combine authentic New York Times articles with synthetic versions generated by multiple state-of-the-art LLMs including Gemma-2-9b, Mistral-7B, Qwen-2-72B, LLaMA-8B, Yi-Large, and GPT-4-o. The dataset provides original article abstracts as prompts, full human-authored narratives. We establish baseline results for two key tasks: distinguishing human-written from AI-generated text, achieving an accuracy of 58.35\%, and attributing AI texts to their generating models with an accuracy of 8.92\%. By bridging real-world journalistic content with modern generative models, the dataset aims to catalyze the development of robust detection and attribution methods, fostering trust and transparency in the era of generative AI. Our dataset is available at: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Text_Dataset
title A Comprehensive Dataset for Human vs. AI Generated Text Detection
topic Computation and Language
url https://arxiv.org/abs/2510.22874