AI evaluation may bias perceptions: The importance of context in interpreting academic writing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Shang, Yao, Randol
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916047298232320
author Wu, Shang
Yao, Randol
author_facet Wu, Shang
Yao, Randol
contents This paper examines how estimates of AI use in scientific writing can be biased when evaluation methods ignore contextual differences across countries and fields. Using large-scale data on journal publications from Dimensions, we construct AI-likeness benchmarks based on differences between human-written and LLM-rephrased abstracts. We show that a pooled benchmark may confound pre-existing stylistic variation with AI-generated text, producing substantial distortions across country-field groups even in pre-LLM publications. In contrast, country-field-specific benchmarks attenuate such distortions and provide a more credible baseline for comparison. Applying these methods to publications in 2025 reveals that the pooled benchmark systematically overestimates AI use in certain countries and fields while underestimating it in others. These findings highlight the importance of context-aware measurement for accurate and equitable evaluation of AI use in science.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26662
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AI evaluation may bias perceptions: The importance of context in interpreting academic writing
Wu, Shang
Yao, Randol
Computation and Language
Artificial Intelligence
General Economics
Economics
This paper examines how estimates of AI use in scientific writing can be biased when evaluation methods ignore contextual differences across countries and fields. Using large-scale data on journal publications from Dimensions, we construct AI-likeness benchmarks based on differences between human-written and LLM-rephrased abstracts. We show that a pooled benchmark may confound pre-existing stylistic variation with AI-generated text, producing substantial distortions across country-field groups even in pre-LLM publications. In contrast, country-field-specific benchmarks attenuate such distortions and provide a more credible baseline for comparison. Applying these methods to publications in 2025 reveals that the pooled benchmark systematically overestimates AI use in certain countries and fields while underestimating it in others. These findings highlight the importance of context-aware measurement for accurate and equitable evaluation of AI use in science.
title AI evaluation may bias perceptions: The importance of context in interpreting academic writing
topic Computation and Language
Artificial Intelligence
General Economics
Economics
url https://arxiv.org/abs/2605.26662