Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hans, Abhimanyu, Schwarzschild, Avi, Cherepanova, Valeriia, Kazemi, Hamid, Saha, Aniruddha, Goldblum, Micah, Geiping, Jonas, Goldstein, Tom
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910647258710016
author Hans, Abhimanyu
Schwarzschild, Avi
Cherepanova, Valeriia
Kazemi, Hamid
Saha, Aniruddha
Goldblum, Micah
Geiping, Jonas
Goldstein, Tom
author_facet Hans, Abhimanyu
Schwarzschild, Avi
Cherepanova, Valeriia
Kazemi, Hamid
Saha, Aniruddha
Goldblum, Micah
Geiping, Jonas
Goldstein, Tom
contents Detecting text generated by modern large language models is thought to be hard, as both LLMs and humans can exhibit a wide range of complex behaviors. However, we find that a score based on contrasting two closely related language models is highly accurate at separating human-generated and machine-generated text. Based on this mechanism, we propose a novel LLM detector that only requires simple calculations using a pair of pre-trained LLMs. The method, called Binoculars, achieves state-of-the-art accuracy without any training data. It is capable of spotting machine text from a range of modern LLMs without any model-specific modifications. We comprehensively evaluate Binoculars on a number of text sources and in varied situations. Over a wide range of document types, Binoculars detects over 90% of generated samples from ChatGPT (and other LLMs) at a false positive rate of 0.01%, despite not being trained on any ChatGPT data.
format Preprint
id arxiv_https___arxiv_org_abs_2401_12070
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
Hans, Abhimanyu
Schwarzschild, Avi
Cherepanova, Valeriia
Kazemi, Hamid
Saha, Aniruddha
Goldblum, Micah
Geiping, Jonas
Goldstein, Tom
Computation and Language
Artificial Intelligence
Machine Learning
Detecting text generated by modern large language models is thought to be hard, as both LLMs and humans can exhibit a wide range of complex behaviors. However, we find that a score based on contrasting two closely related language models is highly accurate at separating human-generated and machine-generated text. Based on this mechanism, we propose a novel LLM detector that only requires simple calculations using a pair of pre-trained LLMs. The method, called Binoculars, achieves state-of-the-art accuracy without any training data. It is capable of spotting machine text from a range of modern LLMs without any model-specific modifications. We comprehensively evaluate Binoculars on a number of text sources and in varied situations. Over a wide range of document types, Binoculars detects over 90% of generated samples from ChatGPT (and other LLMs) at a false positive rate of 0.01%, despite not being trained on any ChatGPT data.
title Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2401.12070