Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Runheng, Huang, Heyan, Xiao, Xingchen, Wu, Zhijing
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918463283396608
author Liu, Runheng
Huang, Heyan
Xiao, Xingchen
Wu, Zhijing
author_facet Liu, Runheng
Huang, Heyan
Xiao, Xingchen
Wu, Zhijing
contents Large language models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their ability to generate human-like text has raised concerns about potential misuse. This underscores the need for reliable and effective methods to detect LLM-generated text. In this paper, we propose IRM, a novel zero-shot approach that leverages Implicit Reward Models for LLM-generated text detection. Such implicit reward models can be derived from publicly available instruction-tuned and base models. Previous reward-based method relies on preference construction and task-specific fine-tuning. In comparison, IRM requires neither preference collection nor additional training. We evaluate IRM on the DetectRL benchmark and demonstrate that IRM can achieve superior detection performance, outperforms existing zero-shot and supervised methods in LLM-generated text detection.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21223
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
Liu, Runheng
Huang, Heyan
Xiao, Xingchen
Wu, Zhijing
Computation and Language
Artificial Intelligence
Large language models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their ability to generate human-like text has raised concerns about potential misuse. This underscores the need for reliable and effective methods to detect LLM-generated text. In this paper, we propose IRM, a novel zero-shot approach that leverages Implicit Reward Models for LLM-generated text detection. Such implicit reward models can be derived from publicly available instruction-tuned and base models. Previous reward-based method relies on preference construction and task-specific fine-tuning. In comparison, IRM requires neither preference collection nor additional training. We evaluate IRM on the DetectRL benchmark and demonstrate that IRM can achieve superior detection performance, outperforms existing zero-shot and supervised methods in LLM-generated text detection.
title Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.21223