Technical Report on the Pangram AI-Generated Text Classifier

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Emi, Bradley, Spero, Max
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909271353982976
author Emi, Bradley
Spero, Max
author_facet Emi, Bradley
Spero, Max
contents We present Pangram Text, a transformer-based neural network trained to distinguish text written by large language models from text written by humans. Pangram Text outperforms zero-shot methods such as DetectGPT as well as leading commercial AI detection tools with over 38 times lower error rates on a comprehensive benchmark comprised of 10 text domains (student writing, creative writing, scientific writing, books, encyclopedias, news, email, scientific papers, short-form Q&A) and 8 open- and closed-source large language models. We propose a training algorithm, hard negative mining with synthetic mirrors, that enables our classifier to achieve orders of magnitude lower false positive rates on high-data domains such as reviews. Finally, we show that Pangram Text is not biased against nonnative English speakers and generalizes to domains and models unseen during training.
format Preprint
id arxiv_https___arxiv_org_abs_2402_14873
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Technical Report on the Pangram AI-Generated Text Classifier
Emi, Bradley
Spero, Max
Computation and Language
Artificial Intelligence
68T50
I.2.7
We present Pangram Text, a transformer-based neural network trained to distinguish text written by large language models from text written by humans. Pangram Text outperforms zero-shot methods such as DetectGPT as well as leading commercial AI detection tools with over 38 times lower error rates on a comprehensive benchmark comprised of 10 text domains (student writing, creative writing, scientific writing, books, encyclopedias, news, email, scientific papers, short-form Q&A) and 8 open- and closed-source large language models. We propose a training algorithm, hard negative mining with synthetic mirrors, that enables our classifier to achieve orders of magnitude lower false positive rates on high-data domains such as reviews. Finally, we show that Pangram Text is not biased against nonnative English speakers and generalizes to domains and models unseen during training.
title Technical Report on the Pangram AI-Generated Text Classifier
topic Computation and Language
Artificial Intelligence
68T50
I.2.7
url https://arxiv.org/abs/2402.14873