A Comprehensive Dataset for Human vs. AI Generated Text Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914598503841792 |
|---|---|
| author | Roy, Rajarshi Singh, Gurpreet Aziz, Ashhar Bajpai, Shashwat Imanpour, Nasrin Biswas, Shwetangshu Wanaskar, Kapil Patwa, Parth Ghosh, Subhankar Dixit, Shreyas Pal, Nilesh Ranjan Rawte, Vipula Garimella, Ritvik Jena, Gaytri Das, Amitava Sheth, Amit Sharma, Vasu Reganti, Aishwarya Naresh Jain, Vinija Chadha, Aman |
| author_facet | Roy, Rajarshi Singh, Gurpreet Aziz, Ashhar Bajpai, Shashwat Imanpour, Nasrin Biswas, Shwetangshu Wanaskar, Kapil Patwa, Parth Ghosh, Subhankar Dixit, Shreyas Pal, Nilesh Ranjan Rawte, Vipula Garimella, Ritvik Jena, Gaytri Das, Amitava Sheth, Amit Sharma, Vasu Reganti, Aishwarya Naresh Jain, Vinija Chadha, Aman |
| contents | The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authenticity, misinformation, and trustworthiness. Addressing the challenge of reliably detecting AI-generated text and attributing it to specific models requires large-scale, diverse, and well-annotated datasets. In this work, we present a comprehensive dataset comprising over 73,193 text samples that combine authentic New York Times articles with synthetic versions generated by multiple state-of-the-art LLMs including Gemma-2-9b, Mistral-7B, Qwen-2-72B, LLaMA-8B, Yi-Large, and GPT-4-o. The dataset provides original article abstracts as prompts, full human-authored narratives. We establish baseline results for two key tasks: distinguishing human-written from AI-generated text, achieving an accuracy of 58.35\%, and attributing AI texts to their generating models with an accuracy of 8.92\%. By bridging real-world journalistic content with modern generative models, the dataset aims to catalyze the development of robust detection and attribution methods, fostering trust and transparency in the era of generative AI. Our dataset is available at: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Text_Dataset |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_22874 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Comprehensive Dataset for Human vs. AI Generated Text Detection Roy, Rajarshi Singh, Gurpreet Aziz, Ashhar Bajpai, Shashwat Imanpour, Nasrin Biswas, Shwetangshu Wanaskar, Kapil Patwa, Parth Ghosh, Subhankar Dixit, Shreyas Pal, Nilesh Ranjan Rawte, Vipula Garimella, Ritvik Jena, Gaytri Das, Amitava Sheth, Amit Sharma, Vasu Reganti, Aishwarya Naresh Jain, Vinija Chadha, Aman Computation and Language The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authenticity, misinformation, and trustworthiness. Addressing the challenge of reliably detecting AI-generated text and attributing it to specific models requires large-scale, diverse, and well-annotated datasets. In this work, we present a comprehensive dataset comprising over 73,193 text samples that combine authentic New York Times articles with synthetic versions generated by multiple state-of-the-art LLMs including Gemma-2-9b, Mistral-7B, Qwen-2-72B, LLaMA-8B, Yi-Large, and GPT-4-o. The dataset provides original article abstracts as prompts, full human-authored narratives. We establish baseline results for two key tasks: distinguishing human-written from AI-generated text, achieving an accuracy of 58.35\%, and attributing AI texts to their generating models with an accuracy of 8.92\%. By bridging real-world journalistic content with modern generative models, the dataset aims to catalyze the development of robust detection and attribution methods, fostering trust and transparency in the era of generative AI. Our dataset is available at: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Text_Dataset |
| title | A Comprehensive Dataset for Human vs. AI Generated Text Detection |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2510.22874 |