Raidar: geneRative AI Detection viA Rewriting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Chengzhi, Vondrick, Carl, Wang, Hao, Yang, Junfeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929313061797888
author Mao, Chengzhi
Vondrick, Carl
Wang, Hao
Yang, Junfeng
author_facet Mao, Chengzhi
Vondrick, Carl
Wang, Hao
Yang, Junfeng
contents We find that large language models (LLMs) are more likely to modify human-written text than AI-generated text when tasked with rewriting. This tendency arises because LLMs often perceive AI-generated text as high-quality, leading to fewer modifications. We introduce a method to detect AI-generated content by prompting LLMs to rewrite text and calculating the editing distance of the output. We dubbed our geneRative AI Detection viA Rewriting method Raidar. Raidar significantly improves the F1 detection scores of existing AI content detection models -- both academic and commercial -- across various domains, including News, creative writing, student essays, code, Yelp reviews, and arXiv papers, with gains of up to 29 points. Operating solely on word symbols without high-dimensional features, our method is compatible with black box LLMs, and is inherently robust on new content. Our results illustrate the unique imprint of machine-generated text through the lens of the machines themselves.
format Preprint
id arxiv_https___arxiv_org_abs_2401_12970
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Raidar: geneRative AI Detection viA Rewriting
Mao, Chengzhi
Vondrick, Carl
Wang, Hao
Yang, Junfeng
Computation and Language
We find that large language models (LLMs) are more likely to modify human-written text than AI-generated text when tasked with rewriting. This tendency arises because LLMs often perceive AI-generated text as high-quality, leading to fewer modifications. We introduce a method to detect AI-generated content by prompting LLMs to rewrite text and calculating the editing distance of the output. We dubbed our geneRative AI Detection viA Rewriting method Raidar. Raidar significantly improves the F1 detection scores of existing AI content detection models -- both academic and commercial -- across various domains, including News, creative writing, student essays, code, Yelp reviews, and arXiv papers, with gains of up to 29 points. Operating solely on word symbols without high-dimensional features, our method is compatible with black box LLMs, and is inherently robust on new content. Our results illustrate the unique imprint of machine-generated text through the lens of the machines themselves.
title Raidar: geneRative AI Detection viA Rewriting
topic Computation and Language
url https://arxiv.org/abs/2401.12970