EAGLE: A Domain Generalization Framework for AI-generated Text Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhattacharjee, Amrita, Moraffah, Raha, Garland, Joshua, Liu, Huan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929286600982528
author Bhattacharjee, Amrita
Moraffah, Raha
Garland, Joshua
Liu, Huan
author_facet Bhattacharjee, Amrita
Moraffah, Raha
Garland, Joshua
Liu, Huan
contents With the advancement in capabilities of Large Language Models (LLMs), one major step in the responsible and safe use of such LLMs is to be able to detect text generated by these models. While supervised AI-generated text detectors perform well on text generated by older LLMs, with the frequent release of new LLMs, building supervised detectors for identifying text from such new models would require new labeled training data, which is infeasible in practice. In this work, we tackle this problem and propose a domain generalization framework for the detection of AI-generated text from unseen target generators. Our proposed framework, EAGLE, leverages the labeled data that is available so far from older language models and learns features invariant across these generators, in order to detect text generated by an unknown target generator. EAGLE learns such domain-invariant features by combining the representational power of self-supervised contrastive learning with domain adversarial training. Through our experiments we demonstrate how EAGLE effectively achieves impressive performance in detecting text generated by unseen target generators, including recent state-of-the-art ones such as GPT-4 and Claude, reaching detection scores of within 4.7% of a fully supervised detector.
format Preprint
id arxiv_https___arxiv_org_abs_2403_15690
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EAGLE: A Domain Generalization Framework for AI-generated Text Detection
Bhattacharjee, Amrita
Moraffah, Raha
Garland, Joshua
Liu, Huan
Computation and Language
Artificial Intelligence
Machine Learning
With the advancement in capabilities of Large Language Models (LLMs), one major step in the responsible and safe use of such LLMs is to be able to detect text generated by these models. While supervised AI-generated text detectors perform well on text generated by older LLMs, with the frequent release of new LLMs, building supervised detectors for identifying text from such new models would require new labeled training data, which is infeasible in practice. In this work, we tackle this problem and propose a domain generalization framework for the detection of AI-generated text from unseen target generators. Our proposed framework, EAGLE, leverages the labeled data that is available so far from older language models and learns features invariant across these generators, in order to detect text generated by an unknown target generator. EAGLE learns such domain-invariant features by combining the representational power of self-supervised contrastive learning with domain adversarial training. Through our experiments we demonstrate how EAGLE effectively achieves impressive performance in detecting text generated by unseen target generators, including recent state-of-the-art ones such as GPT-4 and Claude, reaching detection scores of within 4.7% of a fully supervised detector.
title EAGLE: A Domain Generalization Framework for AI-generated Text Detection
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2403.15690