Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Saha, Rounak, Juneja, Gurusha, Chaudhuri, Dayita, Sajeevan, Naveeja, Shah, Nihar B, Pruthi, Danish
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914412244238336
author Saha, Rounak
Juneja, Gurusha
Chaudhuri, Dayita
Sajeevan, Naveeja
Shah, Nihar B
Pruthi, Danish
author_facet Saha, Rounak
Juneja, Gurusha
Chaudhuri, Dayita
Sajeevan, Naveeja
Shah, Nihar B
Pruthi, Danish
contents A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction of otherwise human-written reviews. But, are these policies enforceable? To answer this question, we assemble a dataset of peer reviews simulating multiple levels of human-AI collaboration, and evaluate five state-of-the-art detectors, including two commercial systems. Our analysis shows that all detectors misclassify a non-trivial fraction of LLM-polished reviews as AI-generated, thereby risking false accusations of academic misconduct. We further investigate whether peer-review-specific signals, including access to the paper manuscript and the constrained domain of scientific writing, can be leveraged to improve detection. While incorporating such signals yields measurable gains in some settings, we identify limitations in each approach and find that none meets the accuracy standards required for identifying AI use in peer reviews. Importantly, our results suggest that recent public estimates of AI use in peer reviews through the use of AI-text detectors should be interpreted with caution, as current detectors misclassify mixed reviews (collaborative human-AI outputs) as fully AI generated, potentially overstating the extent of policy violations.
format Preprint
id arxiv_https___arxiv_org_abs_2603_20450
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
Saha, Rounak
Juneja, Gurusha
Chaudhuri, Dayita
Sajeevan, Naveeja
Shah, Nihar B
Pruthi, Danish
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction of otherwise human-written reviews. But, are these policies enforceable? To answer this question, we assemble a dataset of peer reviews simulating multiple levels of human-AI collaboration, and evaluate five state-of-the-art detectors, including two commercial systems. Our analysis shows that all detectors misclassify a non-trivial fraction of LLM-polished reviews as AI-generated, thereby risking false accusations of academic misconduct. We further investigate whether peer-review-specific signals, including access to the paper manuscript and the constrained domain of scientific writing, can be leveraged to improve detection. While incorporating such signals yields measurable gains in some settings, we identify limitations in each approach and find that none meets the accuracy standards required for identifying AI use in peer reviews. Importantly, our results suggest that recent public estimates of AI use in peer reviews through the use of AI-text detectors should be interpreted with caution, as current detectors misclassify mixed reviews (collaborative human-AI outputs) as fully AI generated, potentially overstating the extent of policy violations.
title Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2603.20450