Are We Truly Innovating? A Qualitative and Quantitative Study of Originality in AI Research Papers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mostafa, Abeer, Nguyen, Thi Huyen, Ahmadi, Zahra
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913167942090752
author Mostafa, Abeer
Nguyen, Thi Huyen
Ahmadi, Zahra
author_facet Mostafa, Abeer
Nguyen, Thi Huyen
Ahmadi, Zahra
contents Assessing originality in AI research is arguably the most consequential yet least reliable step in peer review. Reviewer judgments of originality remain opaque, inconsistent, and dependent on comparisons to prior work that are often incomplete. In this paper, we present a large-scale, data-driven qualitative and quantitative analysis of research originality based on over 100,000 peer-review reports from leading AI venues, spanning a period of rapid growth in the field. Leveraging structured, semantically retrieved prior work and signals embedded in expert reviewer assessments, we systematically characterize how originality is perceived in practice and identify the key dimensions that most strongly influence novelty judgments. Our analysis yields a fine-grained, evidence-based framework that equips both authors and reviewers with actionable insights into how originality is evaluated. In addition, we evaluate the reliability of current large language model (LLM) agents in assessing originality. We find that these models tend to systematically overestimate novelty and struggle to detect conceptual plagiarism, particularly in the presence of paraphrasing. We release our dataset, trained models, and code at: https://anonymous.4open.science/r/Novelty-Reviewer-365C/.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06054
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Are We Truly Innovating? A Qualitative and Quantitative Study of Originality in AI Research Papers
Mostafa, Abeer
Nguyen, Thi Huyen
Ahmadi, Zahra
Computation and Language
Assessing originality in AI research is arguably the most consequential yet least reliable step in peer review. Reviewer judgments of originality remain opaque, inconsistent, and dependent on comparisons to prior work that are often incomplete. In this paper, we present a large-scale, data-driven qualitative and quantitative analysis of research originality based on over 100,000 peer-review reports from leading AI venues, spanning a period of rapid growth in the field. Leveraging structured, semantically retrieved prior work and signals embedded in expert reviewer assessments, we systematically characterize how originality is perceived in practice and identify the key dimensions that most strongly influence novelty judgments. Our analysis yields a fine-grained, evidence-based framework that equips both authors and reviewers with actionable insights into how originality is evaluated. In addition, we evaluate the reliability of current large language model (LLM) agents in assessing originality. We find that these models tend to systematically overestimate novelty and struggle to detect conceptual plagiarism, particularly in the presence of paraphrasing. We release our dataset, trained models, and code at: https://anonymous.4open.science/r/Novelty-Reviewer-365C/.
title Are We Truly Innovating? A Qualitative and Quantitative Study of Originality in AI Research Papers
topic Computation and Language
url https://arxiv.org/abs/2602.06054