Detecting Subtle Differences between Human and Model Languages Using Spectrum of Relative Likelihood

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yang, Wang, Yu, An, Hao, Liu, Zhichen, Li, Yongyuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909341611720704
author Xu, Yang
Wang, Yu
An, Hao
Liu, Zhichen
Li, Yongyuan
author_facet Xu, Yang
Wang, Yu
An, Hao
Liu, Zhichen
Li, Yongyuan
contents Human and model-generated texts can be distinguished by examining the magnitude of likelihood in language. However, it is becoming increasingly difficult as language model's capabilities of generating human-like texts keep evolving. This study provides a new perspective by using the relative likelihood values instead of absolute ones, and extracting useful features from the spectrum-view of likelihood for the human-model text detection task. We propose a detection procedure with two classification methods, supervised and heuristic-based, respectively, which results in competitive performances with previous zero-shot detection methods and a new state-of-the-art on short-text detection. Our method can also reveal subtle differences between human and model languages, which find theoretical roots in psycholinguistics studies. Our code is available at https://github.com/CLCS-SUSTech/FourierGPT
format Preprint
id arxiv_https___arxiv_org_abs_2406_19874
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Detecting Subtle Differences between Human and Model Languages Using Spectrum of Relative Likelihood
Xu, Yang
Wang, Yu
An, Hao
Liu, Zhichen
Li, Yongyuan
Computation and Language
Artificial Intelligence
I.2.7
Human and model-generated texts can be distinguished by examining the magnitude of likelihood in language. However, it is becoming increasingly difficult as language model's capabilities of generating human-like texts keep evolving. This study provides a new perspective by using the relative likelihood values instead of absolute ones, and extracting useful features from the spectrum-view of likelihood for the human-model text detection task. We propose a detection procedure with two classification methods, supervised and heuristic-based, respectively, which results in competitive performances with previous zero-shot detection methods and a new state-of-the-art on short-text detection. Our method can also reveal subtle differences between human and model languages, which find theoretical roots in psycholinguistics studies. Our code is available at https://github.com/CLCS-SUSTech/FourierGPT
title Detecting Subtle Differences between Human and Model Languages Using Spectrum of Relative Likelihood
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2406.19874