Robust and Fine-Grained Detection of AI Generated Texts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kadiyala, Ram Mohan Rao, Pullakhandam, Siddartha, Mehreen, Kanwal, Sharma, Drishti, Gupta, Siddhant, Purbey, Jebish, Srivastava, Ashay, TippaReddy, Subhasya, Bobbili, Arvind Reddy, Chandrashekhar, Suraj Telugara, Adeeb, Modabbir, Vura, Srinadh, Debnath, Suman, Farooq, Hamza
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909713411604480
author Kadiyala, Ram Mohan Rao
Pullakhandam, Siddartha
Mehreen, Kanwal
Sharma, Drishti
Gupta, Siddhant
Purbey, Jebish
Srivastava, Ashay
TippaReddy, Subhasya
Bobbili, Arvind Reddy
Chandrashekhar, Suraj Telugara
Adeeb, Modabbir
Vura, Srinadh
Debnath, Suman
Farooq, Hamza
author_facet Kadiyala, Ram Mohan Rao
Pullakhandam, Siddartha
Mehreen, Kanwal
Sharma, Drishti
Gupta, Siddhant
Purbey, Jebish
Srivastava, Ashay
TippaReddy, Subhasya
Bobbili, Arvind Reddy
Chandrashekhar, Suraj Telugara
Adeeb, Modabbir
Vura, Srinadh
Debnath, Suman
Farooq, Hamza
contents An ideal detection system for machine generated content is supposed to work well on any generator as many more advanced LLMs come into existence day by day. Existing systems often struggle with accurately identifying AI-generated content over shorter texts. Further, not all texts might be entirely authored by a human or LLM, hence we focused more over partial cases i.e human-LLM co-authored texts. Our paper introduces a set of models built for the task of token classification which are trained on an extensive collection of human-machine co-authored texts, which performed well over texts of unseen domains, unseen generators, texts by non-native speakers and those with adversarial inputs. We also introduce a new dataset of over 2.4M such texts mostly co-authored by several popular proprietary LLMs over 23 languages. We also present findings of our models' performance over each texts of each domain and generator. Additional findings include comparison of performance against each adversarial method, length of input texts and characteristics of generated texts compared to the original human authored texts.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11952
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robust and Fine-Grained Detection of AI Generated Texts
Kadiyala, Ram Mohan Rao
Pullakhandam, Siddartha
Mehreen, Kanwal
Sharma, Drishti
Gupta, Siddhant
Purbey, Jebish
Srivastava, Ashay
TippaReddy, Subhasya
Bobbili, Arvind Reddy
Chandrashekhar, Suraj Telugara
Adeeb, Modabbir
Vura, Srinadh
Debnath, Suman
Farooq, Hamza
Computation and Language
Artificial Intelligence
Machine Learning
An ideal detection system for machine generated content is supposed to work well on any generator as many more advanced LLMs come into existence day by day. Existing systems often struggle with accurately identifying AI-generated content over shorter texts. Further, not all texts might be entirely authored by a human or LLM, hence we focused more over partial cases i.e human-LLM co-authored texts. Our paper introduces a set of models built for the task of token classification which are trained on an extensive collection of human-machine co-authored texts, which performed well over texts of unseen domains, unseen generators, texts by non-native speakers and those with adversarial inputs. We also introduce a new dataset of over 2.4M such texts mostly co-authored by several popular proprietary LLMs over 23 languages. We also present findings of our models' performance over each texts of each domain and generator. Additional findings include comparison of performance against each adversarial method, length of input texts and characteristics of generated texts compared to the original human authored texts.
title Robust and Fine-Grained Detection of AI Generated Texts
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.11952