Base Models Look Human To AI Detectors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yixuan Even, Zhong, Ziqian, Raghunathan, Aditi, Fang, Fei, Kolter, J. Zico
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913146104446976
author Xu, Yixuan Even
Zhong, Ziqian
Raghunathan, Aditi
Fang, Fei
Kolter, J. Zico
author_facet Xu, Yixuan Even
Zhong, Ziqian
Raghunathan, Aditi
Fang, Fei
Kolter, J. Zico
contents As AI-generated text enters the real-world at scale, institutions increasingly use commercial AI-text detectors, especially in education and academic-integrity workflows. We report a surprising empirical finding about such systems: when evaluated by GPTZero and Pangram, generated text from base models is often judged overwhelmingly human, whereas text generated by their instruction-tuned counterparts is not. Building on this observation, we propose Humanization by Iterative Paraphrasing (HIP), a detector-agnostic pipeline that minimally fine-tunes a base model into a paraphraser and applies it iteratively. Compared with the baselines we test, HIP yields a stronger trade-off between semantic preservation and detector evasion on commercial detectors. Across Llama-3 and Qwen-3 families, spanning model sizes from 0.6B to 70B, HIP consistently improves detector human-likeness. Our findings suggest that current detectors are tracking artifacts of instruction tuning and local context more than any invariant notion of machine-generated text. This, in turn, calls for detector designs that model these factors more explicitly.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19516
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Base Models Look Human To AI Detectors
Xu, Yixuan Even
Zhong, Ziqian
Raghunathan, Aditi
Fang, Fei
Kolter, J. Zico
Computation and Language
Artificial Intelligence
Machine Learning
As AI-generated text enters the real-world at scale, institutions increasingly use commercial AI-text detectors, especially in education and academic-integrity workflows. We report a surprising empirical finding about such systems: when evaluated by GPTZero and Pangram, generated text from base models is often judged overwhelmingly human, whereas text generated by their instruction-tuned counterparts is not. Building on this observation, we propose Humanization by Iterative Paraphrasing (HIP), a detector-agnostic pipeline that minimally fine-tunes a base model into a paraphraser and applies it iteratively. Compared with the baselines we test, HIP yields a stronger trade-off between semantic preservation and detector evasion on commercial detectors. Across Llama-3 and Qwen-3 families, spanning model sizes from 0.6B to 70B, HIP consistently improves detector human-likeness. Our findings suggest that current detectors are tracking artifacts of instruction tuning and local context more than any invariant notion of machine-generated text. This, in turn, calls for detector designs that model these factors more explicitly.
title Base Models Look Human To AI Detectors
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.19516