Peer Reviews of Peer Reviews: A Randomized Controlled Trial and Other Experiments

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Goldberg, Alexander, Stelmakh, Ivan, Cho, Kyunghyun, Oh, Alice, Agarwal, Alekh, Belgrave, Danielle, Shah, Nihar B.
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913573343592448
author Goldberg, Alexander
Stelmakh, Ivan
Cho, Kyunghyun
Oh, Alice
Agarwal, Alekh
Belgrave, Danielle
Shah, Nihar B.
author_facet Goldberg, Alexander
Stelmakh, Ivan
Cho, Kyunghyun
Oh, Alice
Agarwal, Alekh
Belgrave, Danielle
Shah, Nihar B.
contents Is it possible to reliably evaluate the quality of peer reviews? We study this question driven by two primary motivations -- incentivizing high-quality reviewing using assessed quality of reviews and measuring changes to review quality in experiments. We conduct a large scale study at the NeurIPS 2022 conference, a top-tier conference in machine learning, in which we invited (meta)-reviewers and authors to evaluate reviews given to submitted papers. First, we conduct a RCT to examine bias due to the length of reviews. We generate elongated versions of reviews by adding substantial amounts of non-informative content. Participants in the control group evaluate the original reviews, whereas participants in the experimental group evaluate the artificially lengthened versions. We find that lengthened reviews are scored (statistically significantly) higher quality than the original reviews. In analysis of observational data we find that authors are positively biased towards reviews recommending acceptance of their own papers, even after controlling for confounders of review length, quality, and different numbers of papers per author. We also measure disagreement rates between multiple evaluations of the same review of 28%-32%, which is comparable to that of paper reviewers at NeurIPS. Further, we assess the amount of miscalibration of evaluators of reviews using a linear model of quality scores and find that it is similar to estimates of miscalibration of paper reviewers at NeurIPS. Finally, we estimate the amount of variability in subjective opinions around how to map individual criteria to overall scores of review quality and find that it is roughly the same as that in the review of papers. Our results suggest that the various problems that exist in reviews of papers -- inconsistency, bias towards irrelevant factors, miscalibration, subjectivity -- also arise in reviewing of reviews.
format Preprint
id arxiv_https___arxiv_org_abs_2311_09497
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Peer Reviews of Peer Reviews: A Randomized Controlled Trial and Other Experiments
Goldberg, Alexander
Stelmakh, Ivan
Cho, Kyunghyun
Oh, Alice
Agarwal, Alekh
Belgrave, Danielle
Shah, Nihar B.
Digital Libraries
Computer Science and Game Theory
Is it possible to reliably evaluate the quality of peer reviews? We study this question driven by two primary motivations -- incentivizing high-quality reviewing using assessed quality of reviews and measuring changes to review quality in experiments. We conduct a large scale study at the NeurIPS 2022 conference, a top-tier conference in machine learning, in which we invited (meta)-reviewers and authors to evaluate reviews given to submitted papers. First, we conduct a RCT to examine bias due to the length of reviews. We generate elongated versions of reviews by adding substantial amounts of non-informative content. Participants in the control group evaluate the original reviews, whereas participants in the experimental group evaluate the artificially lengthened versions. We find that lengthened reviews are scored (statistically significantly) higher quality than the original reviews. In analysis of observational data we find that authors are positively biased towards reviews recommending acceptance of their own papers, even after controlling for confounders of review length, quality, and different numbers of papers per author. We also measure disagreement rates between multiple evaluations of the same review of 28%-32%, which is comparable to that of paper reviewers at NeurIPS. Further, we assess the amount of miscalibration of evaluators of reviews using a linear model of quality scores and find that it is similar to estimates of miscalibration of paper reviewers at NeurIPS. Finally, we estimate the amount of variability in subjective opinions around how to map individual criteria to overall scores of review quality and find that it is roughly the same as that in the review of papers. Our results suggest that the various problems that exist in reviews of papers -- inconsistency, bias towards irrelevant factors, miscalibration, subjectivity -- also arise in reviewing of reviews.
title Peer Reviews of Peer Reviews: A Randomized Controlled Trial and Other Experiments
topic Digital Libraries
Computer Science and Game Theory
url https://arxiv.org/abs/2311.09497