Trace Is In Sentences: Unbiased Lightweight ChatGPT-Generated Text Detector

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mu, Mo, Lei, Dianqiao, Li, Chang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909802051928064
author Mu, Mo
Lei, Dianqiao
Li, Chang
author_facet Mu, Mo
Lei, Dianqiao
Li, Chang
contents The widespread adoption of ChatGPT has raised concerns about its misuse, highlighting the need for robust detection of AI-generated text. Current word-level detectors are vulnerable to paraphrasing or simple prompts (PSP), suffer from biases induced by ChatGPT's word-level patterns (CWP) and training data content, degrade on modified text, and often require large models or online LLM interaction. To tackle these issues, we introduce a novel task to detect both original and PSP-modified AI-generated texts, and propose a lightweight framework that classifies texts based on their internal structure, which remains invariant under word-level changes. Our approach encodes sentence embeddings from pre-trained language models and models their relationships via attention. We employ contrastive learning to mitigate embedding biases from autoregressive generation and incorporate a causal graph with counterfactual methods to isolate structural features from topic-related biases. Experiments on two curated datasets, including abstract comparisons and revised life FAQs, validate the effectiveness of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18535
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Trace Is In Sentences: Unbiased Lightweight ChatGPT-Generated Text Detector
Mu, Mo
Lei, Dianqiao
Li, Chang
Computation and Language
Signal Processing
The widespread adoption of ChatGPT has raised concerns about its misuse, highlighting the need for robust detection of AI-generated text. Current word-level detectors are vulnerable to paraphrasing or simple prompts (PSP), suffer from biases induced by ChatGPT's word-level patterns (CWP) and training data content, degrade on modified text, and often require large models or online LLM interaction. To tackle these issues, we introduce a novel task to detect both original and PSP-modified AI-generated texts, and propose a lightweight framework that classifies texts based on their internal structure, which remains invariant under word-level changes. Our approach encodes sentence embeddings from pre-trained language models and models their relationships via attention. We employ contrastive learning to mitigate embedding biases from autoregressive generation and incorporate a causal graph with counterfactual methods to isolate structural features from topic-related biases. Experiments on two curated datasets, including abstract comparisons and revised life FAQs, validate the effectiveness of our method.
title Trace Is In Sentences: Unbiased Lightweight ChatGPT-Generated Text Detector
topic Computation and Language
Signal Processing
url https://arxiv.org/abs/2509.18535