Evaluating Research Quality with Large Language Models: An Analysis of ChatGPT's Effectiveness with Different Settings and Inputs
Fuente:
arXiv
Saved in:
| Main Author: | Thelwall, Mike |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can ChatGPT evaluate research quality?
by: Thelwall, Mike
Published: (2024)
by: Thelwall, Mike
Published: (2024)
Can Smaller Large Language Models Evaluate Research Quality?
by: Thelwall, Mike
Published: (2025)
by: Thelwall, Mike
Published: (2025)
Evaluating the quality of published medical research with ChatGPT
by: Thelwall, Mike, et al.
Published: (2024)
by: Thelwall, Mike, et al.
Published: (2024)
Assessing the societal influence of academic research with ChatGPT: Impact case study evaluations
by: Kousha, Kayvan, et al.
Published: (2024)
by: Kousha, Kayvan, et al.
Published: (2024)
Implicit and Explicit Research Quality Score Probabilities from ChatGPT
by: Thelwall, Mike, et al.
Published: (2025)
by: Thelwall, Mike, et al.
Published: (2025)
A Global South Strategy for Evaluating Research Value with ChatGPT
by: Nunkoo, Robin, et al.
Published: (2025)
by: Nunkoo, Robin, et al.
Published: (2025)
Can Small and Reasoning Large Language Models Score Journal Articles for Research Quality and Do Averaging and Few-shot Help?
by: Thelwall, Mike, et al.
Published: (2025)
by: Thelwall, Mike, et al.
Published: (2025)
Research evaluation with ChatGPT: Is it age, country, length, or field biased?
by: Thelwall, Mike, et al.
Published: (2024)
by: Thelwall, Mike, et al.
Published: (2024)
Journal Quality Factors from ChatGPT: More meaningful than Impact Factors?
by: Thelwall, Mike, et al.
Published: (2024)
by: Thelwall, Mike, et al.
Published: (2024)
In which fields do ChatGPT scores align better than citations with research quality?
by: Thelwall, Mike
Published: (2025)
by: Thelwall, Mike
Published: (2025)
Which stylistic features fool ChatGPT research evaluations?
by: Kousha, Kayvan, et al.
Published: (2026)
by: Kousha, Kayvan, et al.
Published: (2026)
Estimating the quality of academic books from their descriptions with ChatGPT
by: Thelwall, Mike, et al.
Published: (2025)
by: Thelwall, Mike, et al.
Published: (2025)
Evaluating the Predictive Capacity of ChatGPT for Academic Peer Review Outcomes Across Multiple Platforms
by: Thelwall, Mike, et al.
Published: (2024)
by: Thelwall, Mike, et al.
Published: (2024)
In which fields can ChatGPT detect journal article quality? An evaluation of REF2021 results
by: Thelwall, Mike, et al.
Published: (2024)
by: Thelwall, Mike, et al.
Published: (2024)
Can ChatGPT evaluate research environments? Evidence from REF2021
by: Kousha, Kayvan, et al.
Published: (2025)
by: Kousha, Kayvan, et al.
Published: (2025)
Can ChatGPT be a good follower of academic paradigms? Research quality evaluations in conflicting areas of sociology
by: Thelwall, Mike, et al.
Published: (2025)
by: Thelwall, Mike, et al.
Published: (2025)
How much are LLMs changing the language of academic papers after ChatGPT? A multi-database and full text analysis
by: Kousha, Kayvan, et al.
Published: (2025)
by: Kousha, Kayvan, et al.
Published: (2025)
Research quality evaluation by AI in the era of Large Language Models: Advantages, disadvantages, and systemic effects
by: Thelwall, Mike
Published: (2025)
by: Thelwall, Mike
Published: (2025)
Can Large Language Models Evaluate Grant Proposal Quality? Revisiting the Wennerås and Wold Peer Review Data
by: Sandström, Ulf, et al.
Published: (2026)
by: Sandström, Ulf, et al.
Published: (2026)
Prompt perturbation and fraction facilitation sometimes strengthen Large Language Model scores
by: Thelwall, Mike
Published: (2025)
by: Thelwall, Mike
Published: (2025)
Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses
by: Jacques, Erin, et al.
Published: (2026)
by: Jacques, Erin, et al.
Published: (2026)
Do Large Language Models know Which Published Articles have been Retracted?
by: Thelwall, Mike
Published: (2026)
by: Thelwall, Mike
Published: (2026)
Is OpenAlex Suitable for Research Quality Evaluation and Which Citation Indicator is Best?
by: Thelwall, Mike, et al.
Published: (2025)
by: Thelwall, Mike, et al.
Published: (2025)
Quantitative Methods in Research Evaluation Citation Indicators, Altmetrics, and Artificial Intelligence
by: Thelwall, Mike
Published: (2024)
by: Thelwall, Mike
Published: (2024)
FAIR GPT: A virtual consultant for research data management in ChatGPT
by: Shigapov, Renat, et al.
Published: (2024)
by: Shigapov, Renat, et al.
Published: (2024)
Is ChatGPT Transforming Academics' Writing Style?
by: Geng, Mingmeng, et al.
Published: (2024)
by: Geng, Mingmeng, et al.
Published: (2024)
An Experiment with the Use of ChatGPT for LCSH Subject Assignment on Electronic Theses and Dissertations
by: Chow, Eric H. C., et al.
Published: (2024)
by: Chow, Eric H. C., et al.
Published: (2024)
Large Language Models for Departmental Expert Review Quality Scores
by: Langfeldt, Liv, et al.
Published: (2026)
by: Langfeldt, Liv, et al.
Published: (2026)
From ChatGPT, DALL-E 3 to Sora: How has Generative AI Changed Digital Humanities Research and Services?
by: Liu, Jiangfeng, et al.
Published: (2024)
by: Liu, Jiangfeng, et al.
Published: (2024)
Are ChatGPT and Other Similar Systems the Modern Lernaean Hydras of AI?
by: Ioannidis, Dimitrios, et al.
Published: (2023)
by: Ioannidis, Dimitrios, et al.
Published: (2023)
Will AI be overconfident about academic research findings when reliant on abstracts? (v1)
by: Thelwall, Mike
Published: (2026)
by: Thelwall, Mike
Published: (2026)
Designing large language model prompts to extract scores from messy text: A shared dataset and challenge
by: Thelwall, Mike
Published: (2026)
by: Thelwall, Mike
Published: (2026)
Can social media provide early warning of retraction? Evidence from critical tweets identified by human annotation and large language models
by: Zheng, Er-Te, et al.
Published: (2024)
by: Zheng, Er-Te, et al.
Published: (2024)
Have LLM-associated terms increased in article full texts in all fields?
by: Thelwall, Mike, et al.
Published: (2026)
by: Thelwall, Mike, et al.
Published: (2026)
Bridging the Evaluation Gap: Leveraging Large Language Models for Topic Model Evaluation
by: Tan, Zhiyin, et al.
Published: (2025)
by: Tan, Zhiyin, et al.
Published: (2025)
ChatGPT "contamination": estimating the prevalence of LLMs in the scholarly literature
by: Gray, Andrew
Published: (2024)
by: Gray, Andrew
Published: (2024)
LLAssist: Simple Tools for Automating Literature Review Using Large Language Models
by: Haryanto, Christoforus Yoga
Published: (2024)
by: Haryanto, Christoforus Yoga
Published: (2024)
Toward Purpose-oriented Topic Model Evaluation enabled by Large Language Models
by: Tan, Zhiyin, et al.
Published: (2025)
by: Tan, Zhiyin, et al.
Published: (2025)
Matching Game Preferences Through Dialogical Large Language Models: A Perspective
by: Fabre, Renaud, et al.
Published: (2025)
by: Fabre, Renaud, et al.
Published: (2025)
Towards Large Language Models for Lunar Mission Planning and In Situ Resource Utilization
by: Pekala, Michael, et al.
Published: (2025)
by: Pekala, Michael, et al.
Published: (2025)
Similar Items
-
Can ChatGPT evaluate research quality?
by: Thelwall, Mike
Published: (2024) -
Can Smaller Large Language Models Evaluate Research Quality?
by: Thelwall, Mike
Published: (2025) -
Evaluating the quality of published medical research with ChatGPT
by: Thelwall, Mike, et al.
Published: (2024) -
Assessing the societal influence of academic research with ChatGPT: Impact case study evaluations
by: Kousha, Kayvan, et al.
Published: (2024) -
Implicit and Explicit Research Quality Score Probabilities from ChatGPT
by: Thelwall, Mike, et al.
Published: (2025)