Saved in:
Bibliographic Details
Main Authors: Shi, Yuhui, Sheng, Qiang, Cao, Juan, Mi, Hao, Hu, Beizhe, Wang, Danding
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2402.09199
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929476985683968
author Shi, Yuhui
Sheng, Qiang
Cao, Juan
Mi, Hao
Hu, Beizhe
Wang, Danding
author_facet Shi, Yuhui
Sheng, Qiang
Cao, Juan
Mi, Hao
Hu, Beizhe
Wang, Danding
contents With the rapidly increasing application of large language models (LLMs), their abuse has caused many undesirable societal problems such as fake news, academic dishonesty, and information pollution. This makes AI-generated text (AIGT) detection of great importance. Among existing methods, white-box methods are generally superior to black-box methods in terms of performance and generalizability, but they require access to LLMs' internal states and are not applicable to black-box settings. In this paper, we propose to estimate word generation probabilities as pseudo white-box features via multiple re-sampling to help improve AIGT detection under the black-box setting. Specifically, we design POGER, a proxy-guided efficient re-sampling method, which selects a small subset of representative words (e.g., 10 words) for performing multiple re-sampling in black-box AIGT detection. Experiments on datasets containing texts from humans and seven LLMs show that POGER outperforms all baselines in macro F1 under black-box, partial white-box, and out-of-distribution settings and maintains lower re-sampling costs than its existing counterparts.
format Preprint
id arxiv_https___arxiv_org_abs_2402_09199
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Ten Words Only Still Help: Improving Black-Box AI-Generated Text Detection via Proxy-Guided Efficient Re-Sampling
Shi, Yuhui
Sheng, Qiang
Cao, Juan
Mi, Hao
Hu, Beizhe
Wang, Danding
Computation and Language
Artificial Intelligence
Machine Learning
With the rapidly increasing application of large language models (LLMs), their abuse has caused many undesirable societal problems such as fake news, academic dishonesty, and information pollution. This makes AI-generated text (AIGT) detection of great importance. Among existing methods, white-box methods are generally superior to black-box methods in terms of performance and generalizability, but they require access to LLMs' internal states and are not applicable to black-box settings. In this paper, we propose to estimate word generation probabilities as pseudo white-box features via multiple re-sampling to help improve AIGT detection under the black-box setting. Specifically, we design POGER, a proxy-guided efficient re-sampling method, which selects a small subset of representative words (e.g., 10 words) for performing multiple re-sampling in black-box AIGT detection. Experiments on datasets containing texts from humans and seven LLMs show that POGER outperforms all baselines in macro F1 under black-box, partial white-box, and out-of-distribution settings and maintains lower re-sampling costs than its existing counterparts.
title Ten Words Only Still Help: Improving Black-Box AI-Generated Text Detection via Proxy-Guided Efficient Re-Sampling
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2402.09199