| _version_ | 1866901570407366656 |
|---|---|
| author | Campbell, Colin |
| author_facet | Campbell, Colin |
| contents | <p class="MsoNormal">This study evaluated ChatGPT-4o's potential as a scalable tool for formative assessment in English-as-a-foreign-language (EFL) writing instruction in higher education, with data drawn from a South Korean university context. Using a mixed-methods design that combined Item Response Theory (IRT), qualitative feedback analysis, and the Assessment for Learning (AfL) framework, the study compared ChatGPT-4o's holistic essay scoring and qualitative feedback against those of three experienced, IELTS-certified university English instructors. A total of 76 essays produced by 38 non-English-major undergraduates for the IELTS Academic Writing Test (Task 1 and Task 2) were analyzed during the Fall 2024 semester (September–December 2024). ChatGPT-4o and the human raters scored the essays using the IELTS rubric and provided comments on language use, content quality, and organizational structure. IRT-based results indicate that ChatGPT-4o accounted for 87.63% of person variance compared to 76.59% for human raters, with significantly lower residual variance (10.83% vs. 18.58%), suggesting stronger internal scoring consistency. No statistically significant differences in mean scores were found (F = 0.078, p = 0.972; ICC = 0.792; α = 0.937). Qualitative feedback analysis revealed that ChatGPT-4o delivered more balanced and comprehensive feedback with a strong emphasis on content and organizational structure, whereas teacher feedback focused more heavily on surface-level linguistic accuracy. These findings have implications for AI-assisted writing assessment in large-enrollment EFL contexts globally, where scalable, consistent formative feedback remains a persistent challenge.</p> <p><br><br></p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">The datasets supporting the findings of this study are openly available and consist of four files. The primary quantitative file, <em>scores.csv</em>, contains 532 rows recording holistic and sub-criterion scores (Task Achievement, Coherence and Cohesion, Lexical Resource, and Grammatical Range and Accuracy) for each of the 38 student participants across both writing tasks, four raters, and — for ChatGPT-4 — four temporally separated scoring runs. The qualitative dataset, <em>feedback_coding.csv</em>, contains 912 rows representing the full thematic coding of rater feedback, with each row corresponding to one student, one task, one rater, and one of the three feedback categories (Language, Content, Organization). Binary flags for surface-level focus, actionability, and higher-order concern are included for each coded instance. The psychometric output file, <em>irt_parameters.csv</em>, provides the Many-Facet Rasch Model estimates for all four raters, including severity, infit and outfit mean-square statistics, person variance decomposition, and aggregate reliability indices as reported in Tables 1, 3, and 4 of this paper.</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">All variable definitions, value ranges, coding procedures, theme code descriptions, and inter-coder reliability notes are documented in the accompanying <em>codebook.pdf</em>. Student identifiers are fully anonymized. Data were collected under written informed consent.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19656187 |
| institution | Zenodo |
| language | |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | AI-Driven Formative Assessment in EFL Writing: A Comparative Study of ChatGPT-4o and Human Raters Campbell, Colin <p class="MsoNormal">This study evaluated ChatGPT-4o's potential as a scalable tool for formative assessment in English-as-a-foreign-language (EFL) writing instruction in higher education, with data drawn from a South Korean university context. Using a mixed-methods design that combined Item Response Theory (IRT), qualitative feedback analysis, and the Assessment for Learning (AfL) framework, the study compared ChatGPT-4o's holistic essay scoring and qualitative feedback against those of three experienced, IELTS-certified university English instructors. A total of 76 essays produced by 38 non-English-major undergraduates for the IELTS Academic Writing Test (Task 1 and Task 2) were analyzed during the Fall 2024 semester (September–December 2024). ChatGPT-4o and the human raters scored the essays using the IELTS rubric and provided comments on language use, content quality, and organizational structure. IRT-based results indicate that ChatGPT-4o accounted for 87.63% of person variance compared to 76.59% for human raters, with significantly lower residual variance (10.83% vs. 18.58%), suggesting stronger internal scoring consistency. No statistically significant differences in mean scores were found (F = 0.078, p = 0.972; ICC = 0.792; α = 0.937). Qualitative feedback analysis revealed that ChatGPT-4o delivered more balanced and comprehensive feedback with a strong emphasis on content and organizational structure, whereas teacher feedback focused more heavily on surface-level linguistic accuracy. These findings have implications for AI-assisted writing assessment in large-enrollment EFL contexts globally, where scalable, consistent formative feedback remains a persistent challenge.</p> <p><br><br></p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">The datasets supporting the findings of this study are openly available and consist of four files. The primary quantitative file, <em>scores.csv</em>, contains 532 rows recording holistic and sub-criterion scores (Task Achievement, Coherence and Cohesion, Lexical Resource, and Grammatical Range and Accuracy) for each of the 38 student participants across both writing tasks, four raters, and — for ChatGPT-4 — four temporally separated scoring runs. The qualitative dataset, <em>feedback_coding.csv</em>, contains 912 rows representing the full thematic coding of rater feedback, with each row corresponding to one student, one task, one rater, and one of the three feedback categories (Language, Content, Organization). Binary flags for surface-level focus, actionability, and higher-order concern are included for each coded instance. The psychometric output file, <em>irt_parameters.csv</em>, provides the Many-Facet Rasch Model estimates for all four raters, including severity, infit and outfit mean-square statistics, person variance decomposition, and aggregate reliability indices as reported in Tables 1, 3, and 4 of this paper.</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">All variable definitions, value ranges, coding procedures, theme code descriptions, and inter-coder reliability notes are documented in the accompanying <em>codebook.pdf</em>. Student identifiers are fully anonymized. Data were collected under written informed consent.</p> |
| title | AI-Driven Formative Assessment in EFL Writing: A Comparative Study of ChatGPT-4o and Human Raters |
| url | https://doi.org/10.5281/zenodo.19656187 |