Can ChatGPT evaluate research environments? Evidence from REF2021

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kousha, Kayvan, Thelwall, Mike, Gadd, Elizabeth
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909944063721472
author Kousha, Kayvan
Thelwall, Mike
Gadd, Elizabeth
author_facet Kousha, Kayvan
Thelwall, Mike
Gadd, Elizabeth
contents UK academic departments are evaluated partly on the statements that they write about the value of their research environments for the Research Excellence Framework (REF) periodic assessments. These statements mix qualitative narratives and quantitative data, typically requiring time-consuming and difficult expert judgements to assess. This article investigates whether Large Language Models (LLMs) can support the process or validate the results, using the UK REF2021 unit-level environment statements as a test case. Based on prompts mimicking the REF guidelines, ChatGPT 4o-mini scores correlated positively with expert scores in almost all 34 (field-based) Units of Assessment (UoAs). ChatGPT's scores had moderate to strong positive Spearman correlations with REF expert scores in 32 out of 34 UoAs: 14 UoAs above 0.7 and a further 13 between 0.6 and 0.7. Only two UoAs had weak or no significant associations (Classics and Clinical Medicine). From further tests for UoA34, multiple LLMs had significant positive correlations with REF2021 environment scores (all p < .001), with ChatGPT 5 performing best (r=0.81; $ρ$=0.82), followed by ChatGPT-4o-mini (r=0.68; $ρ$=0.67) and Gemini Flash 2.5 (r=0.67; $ρ$=0.69). If LLM-generated scores for environment statements are used in future to help reduce workload, support more consistent interpretation, and complement human review then caution must be exercised because of the potential for biases, inaccuracy in some cases, and unwanted systemic effects. Even the strong correlations found here seem unlikely to be judged close enough to expert scores to fully delegate the assessment task to LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05202
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can ChatGPT evaluate research environments? Evidence from REF2021
Kousha, Kayvan
Thelwall, Mike
Gadd, Elizabeth
Digital Libraries
UK academic departments are evaluated partly on the statements that they write about the value of their research environments for the Research Excellence Framework (REF) periodic assessments. These statements mix qualitative narratives and quantitative data, typically requiring time-consuming and difficult expert judgements to assess. This article investigates whether Large Language Models (LLMs) can support the process or validate the results, using the UK REF2021 unit-level environment statements as a test case. Based on prompts mimicking the REF guidelines, ChatGPT 4o-mini scores correlated positively with expert scores in almost all 34 (field-based) Units of Assessment (UoAs). ChatGPT's scores had moderate to strong positive Spearman correlations with REF expert scores in 32 out of 34 UoAs: 14 UoAs above 0.7 and a further 13 between 0.6 and 0.7. Only two UoAs had weak or no significant associations (Classics and Clinical Medicine). From further tests for UoA34, multiple LLMs had significant positive correlations with REF2021 environment scores (all p < .001), with ChatGPT 5 performing best (r=0.81; $ρ$=0.82), followed by ChatGPT-4o-mini (r=0.68; $ρ$=0.67) and Gemini Flash 2.5 (r=0.67; $ρ$=0.69). If LLM-generated scores for environment statements are used in future to help reduce workload, support more consistent interpretation, and complement human review then caution must be exercised because of the potential for biases, inaccuracy in some cases, and unwanted systemic effects. Even the strong correlations found here seem unlikely to be judged close enough to expert scores to fully delegate the assessment task to LLMs.
title Can ChatGPT evaluate research environments? Evidence from REF2021
topic Digital Libraries
url https://arxiv.org/abs/2512.05202