| _version_ | 1866901570899148800 |
|---|---|
| author | Campbell, Colin |
| author_facet | Campbell, Colin |
| contents | <p><span>This study utilized Item Response Theory (IRT), qualitative feedback analysis, and the Assessment for Learning (AfL) framework to evaluate ChatGPT 4’s potential as a tool for improving English-as-a-foreign-language (EFL) writing assessment in South Korean higher education. The research aimed to assess the reliability of holistic essay scores assigned by ChatGPT 4 compared to those given by experienced university English instructors and to examine the utility of qualitative feedback provided by AI and human raters. A total of 76 essays written by non-English major students for the International English Language Testing System (IELTS) Academic Writing Test for Task 1 and Task 2 were analyzed. ChatGPT 4 and three university English instructors, rated the essays using the IELTS scoring rubric and provided comments on language use, content quality, and organizational structure. Results indicate that ChatGPT 4 and human raters were similarly accurate in scoring quantitative scores. Qualitative feedback analysis found that ChatGPT consistently delivered more balanced and comprehensive feedback, with a strong emphasis on content and organizational structure. By contrast, teacher feedback was often more focused on linguistic accuracy, sometimes overlooking broader aspects of writing.</span></p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_14854639 |
| institution | Zenodo |
| language | |
| publishDate | 2025 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | AI-Driven Formative Assessment in EFL Writing: A Comparative Study of ChatGPT-4 and Human Raters Campbell, Colin <p><span>This study utilized Item Response Theory (IRT), qualitative feedback analysis, and the Assessment for Learning (AfL) framework to evaluate ChatGPT 4’s potential as a tool for improving English-as-a-foreign-language (EFL) writing assessment in South Korean higher education. The research aimed to assess the reliability of holistic essay scores assigned by ChatGPT 4 compared to those given by experienced university English instructors and to examine the utility of qualitative feedback provided by AI and human raters. A total of 76 essays written by non-English major students for the International English Language Testing System (IELTS) Academic Writing Test for Task 1 and Task 2 were analyzed. ChatGPT 4 and three university English instructors, rated the essays using the IELTS scoring rubric and provided comments on language use, content quality, and organizational structure. Results indicate that ChatGPT 4 and human raters were similarly accurate in scoring quantitative scores. Qualitative feedback analysis found that ChatGPT consistently delivered more balanced and comprehensive feedback, with a strong emphasis on content and organizational structure. By contrast, teacher feedback was often more focused on linguistic accuracy, sometimes overlooking broader aspects of writing.</span></p> |
| title | AI-Driven Formative Assessment in EFL Writing: A Comparative Study of ChatGPT-4 and Human Raters |
| url | https://doi.org/10.5281/zenodo.14854639 |