Can We Trust LLM Detectors?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sandhan, Jivnesh, Jaiswal, Harshit, Cheng, Fei, Murawaki, Yugo
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908790011461632
author Sandhan, Jivnesh
Jaiswal, Harshit
Cheng, Fei
Murawaki, Yugo
author_facet Sandhan, Jivnesh
Jaiswal, Harshit
Cheng, Fei
Murawaki, Yugo
contents The rapid adoption of LLMs has increased the need for reliable AI text detection, yet existing detectors often fail outside controlled benchmarks. We systematically evaluate 2 dominant paradigms (training-free and supervised) and show that both are brittle under distribution shift, unseen generators, and simple stylistic perturbations. To address these limitations, we propose a supervised contrastive learning (SCL) framework that learns discriminative style embeddings. Experiments show that while supervised detectors excel in-domain, they degrade sharply out-of-domain, and training-free methods remain highly sensitive to proxy choice. Overall, our results expose fundamental challenges in building domain-agnostic detectors. Our code is available at: https://github.com/HARSHITJAIS14/DetectAI
format Preprint
id arxiv_https___arxiv_org_abs_2601_15301
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Can We Trust LLM Detectors?
Sandhan, Jivnesh
Jaiswal, Harshit
Cheng, Fei
Murawaki, Yugo
Computation and Language
Artificial Intelligence
The rapid adoption of LLMs has increased the need for reliable AI text detection, yet existing detectors often fail outside controlled benchmarks. We systematically evaluate 2 dominant paradigms (training-free and supervised) and show that both are brittle under distribution shift, unseen generators, and simple stylistic perturbations. To address these limitations, we propose a supervised contrastive learning (SCL) framework that learns discriminative style embeddings. Experiments show that while supervised detectors excel in-domain, they degrade sharply out-of-domain, and training-free methods remain highly sensitive to proxy choice. Overall, our results expose fundamental challenges in building domain-agnostic detectors. Our code is available at: https://github.com/HARSHITJAIS14/DetectAI
title Can We Trust LLM Detectors?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.15301