AI-Driven Formative Assessment in EFL Writing: A Comparative Study of ChatGPT-4o and Human Raters

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Author: Campbell, Colin
Format: Recurso digital
Published: Zenodo 2025
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866901570407366656
author Campbell, Colin
author_facet Campbell, Colin
contents <p class="MsoNormal">This study evaluated ChatGPT-4o's potential as a scalable tool for formative assessment in English-as-a-foreign-language (EFL) writing instruction in higher education, with data drawn from a South Korean university context. Using a mixed-methods design that combined Item Response Theory (IRT), qualitative feedback analysis, and the Assessment for Learning (AfL) framework, the study compared ChatGPT-4o's holistic essay scoring and qualitative feedback against those of three experienced, IELTS-certified university English instructors. A total of 76 essays produced by 38 non-English-major undergraduates for the IELTS Academic Writing Test (Task 1 and Task 2) were analyzed during the Fall 2024 semester (September–December 2024). ChatGPT-4o and the human raters scored the essays using the IELTS rubric and provided comments on language use, content quality, and organizational structure. IRT-based results indicate that ChatGPT-4o accounted for 87.63% of person variance compared to 76.59% for human raters, with significantly lower residual variance (10.83% vs. 18.58%), suggesting stronger internal scoring consistency. No statistically significant differences in mean scores were found (F = 0.078, p = 0.972; ICC = 0.792; α = 0.937). Qualitative feedback analysis revealed that ChatGPT-4o delivered more balanced and comprehensive feedback with a strong emphasis on content and organizational structure, whereas teacher feedback focused more heavily on surface-level linguistic accuracy. These findings have implications for AI-assisted writing assessment in large-enrollment EFL contexts globally, where scalable, consistent formative feedback remains a persistent challenge.</p> <p><br><br></p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">The datasets supporting the findings of this study are openly available and consist of four files. The primary quantitative file, <em>scores.csv</em>, contains 532 rows recording holistic and sub-criterion scores (Task Achievement, Coherence and Cohesion, Lexical Resource, and Grammatical Range and Accuracy) for each of the 38 student participants across both writing tasks, four raters, and — for ChatGPT-4 — four temporally separated scoring runs. The qualitative dataset, <em>feedback_coding.csv</em>, contains 912 rows representing the full thematic coding of rater feedback, with each row corresponding to one student, one task, one rater, and one of the three feedback categories (Language, Content, Organization). Binary flags for surface-level focus, actionability, and higher-order concern are included for each coded instance. The psychometric output file, <em>irt_parameters.csv</em>, provides the Many-Facet Rasch Model estimates for all four raters, including severity, infit and outfit mean-square statistics, person variance decomposition, and aggregate reliability indices as reported in Tables 1, 3, and 4 of this paper.</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">All variable definitions, value ranges, coding procedures, theme code descriptions, and inter-coder reliability notes are documented in the accompanying <em>codebook.pdf</em>. Student identifiers are fully anonymized. Data were collected under written informed consent.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19656187
institution Zenodo
language
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle AI-Driven Formative Assessment in EFL Writing: A Comparative Study of ChatGPT-4o and Human Raters
Campbell, Colin
<p class="MsoNormal">This study evaluated ChatGPT-4o's potential as a scalable tool for formative assessment in English-as-a-foreign-language (EFL) writing instruction in higher education, with data drawn from a South Korean university context. Using a mixed-methods design that combined Item Response Theory (IRT), qualitative feedback analysis, and the Assessment for Learning (AfL) framework, the study compared ChatGPT-4o's holistic essay scoring and qualitative feedback against those of three experienced, IELTS-certified university English instructors. A total of 76 essays produced by 38 non-English-major undergraduates for the IELTS Academic Writing Test (Task 1 and Task 2) were analyzed during the Fall 2024 semester (September–December 2024). ChatGPT-4o and the human raters scored the essays using the IELTS rubric and provided comments on language use, content quality, and organizational structure. IRT-based results indicate that ChatGPT-4o accounted for 87.63% of person variance compared to 76.59% for human raters, with significantly lower residual variance (10.83% vs. 18.58%), suggesting stronger internal scoring consistency. No statistically significant differences in mean scores were found (F = 0.078, p = 0.972; ICC = 0.792; α = 0.937). Qualitative feedback analysis revealed that ChatGPT-4o delivered more balanced and comprehensive feedback with a strong emphasis on content and organizational structure, whereas teacher feedback focused more heavily on surface-level linguistic accuracy. These findings have implications for AI-assisted writing assessment in large-enrollment EFL contexts globally, where scalable, consistent formative feedback remains a persistent challenge.</p> <p><br><br></p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">The datasets supporting the findings of this study are openly available and consist of four files. The primary quantitative file, <em>scores.csv</em>, contains 532 rows recording holistic and sub-criterion scores (Task Achievement, Coherence and Cohesion, Lexical Resource, and Grammatical Range and Accuracy) for each of the 38 student participants across both writing tasks, four raters, and — for ChatGPT-4 — four temporally separated scoring runs. The qualitative dataset, <em>feedback_coding.csv</em>, contains 912 rows representing the full thematic coding of rater feedback, with each row corresponding to one student, one task, one rater, and one of the three feedback categories (Language, Content, Organization). Binary flags for surface-level focus, actionability, and higher-order concern are included for each coded instance. The psychometric output file, <em>irt_parameters.csv</em>, provides the Many-Facet Rasch Model estimates for all four raters, including severity, infit and outfit mean-square statistics, person variance decomposition, and aggregate reliability indices as reported in Tables 1, 3, and 4 of this paper.</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">All variable definitions, value ranges, coding procedures, theme code descriptions, and inter-coder reliability notes are documented in the accompanying <em>codebook.pdf</em>. Student identifiers are fully anonymized. Data were collected under written informed consent.</p>
title AI-Driven Formative Assessment in EFL Writing: A Comparative Study of ChatGPT-4o and Human Raters
url https://doi.org/10.5281/zenodo.19656187