AI-Driven Formative Assessment in EFL Writing: A Comparative Study of ChatGPT-4 and Human Raters

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Author: Campbell, Colin
Format: Recurso digital
Published: Zenodo 2025
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866901570899148800
author Campbell, Colin
author_facet Campbell, Colin
contents <p><span>This study utilized Item Response Theory (IRT), qualitative feedback analysis, and the Assessment for Learning (AfL) framework to evaluate ChatGPT 4’s potential as a tool for improving English-as-a-foreign-language (EFL) writing assessment in South Korean higher education. The research aimed to assess the reliability of holistic essay scores assigned by ChatGPT 4 compared to those given by experienced university English instructors and to examine the utility of qualitative feedback provided by AI and human raters. A total of 76 essays written by non-English major students for the International English Language Testing System (IELTS) Academic Writing Test for Task 1 and Task 2 were analyzed. ChatGPT 4 and three university English instructors, rated the essays using the IELTS scoring rubric and provided comments on language use, content quality, and organizational structure. Results indicate that ChatGPT 4 and human raters were similarly accurate in scoring quantitative scores. Qualitative feedback analysis found that ChatGPT consistently delivered more balanced and comprehensive feedback, with a strong emphasis on content and organizational structure. By contrast, teacher feedback was often more focused on linguistic accuracy, sometimes overlooking broader aspects of writing.</span></p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_14854639
institution Zenodo
language
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle AI-Driven Formative Assessment in EFL Writing: A Comparative Study of ChatGPT-4 and Human Raters
Campbell, Colin
<p><span>This study utilized Item Response Theory (IRT), qualitative feedback analysis, and the Assessment for Learning (AfL) framework to evaluate ChatGPT 4’s potential as a tool for improving English-as-a-foreign-language (EFL) writing assessment in South Korean higher education. The research aimed to assess the reliability of holistic essay scores assigned by ChatGPT 4 compared to those given by experienced university English instructors and to examine the utility of qualitative feedback provided by AI and human raters. A total of 76 essays written by non-English major students for the International English Language Testing System (IELTS) Academic Writing Test for Task 1 and Task 2 were analyzed. ChatGPT 4 and three university English instructors, rated the essays using the IELTS scoring rubric and provided comments on language use, content quality, and organizational structure. Results indicate that ChatGPT 4 and human raters were similarly accurate in scoring quantitative scores. Qualitative feedback analysis found that ChatGPT consistently delivered more balanced and comprehensive feedback, with a strong emphasis on content and organizational structure. By contrast, teacher feedback was often more focused on linguistic accuracy, sometimes overlooking broader aspects of writing.</span></p>
title AI-Driven Formative Assessment in EFL Writing: A Comparative Study of ChatGPT-4 and Human Raters
url https://doi.org/10.5281/zenodo.14854639