From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nguyen, Minh Hoang, Pham, Vu Hoang, Huynh, Xuan Thanh, Mai, Phuc Hong, Nguyen, Vinh The, Huynh, Quang Nhut, Nguyen, Huy Tien, Le, Tung
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915839823839232
author Nguyen, Minh Hoang
Pham, Vu Hoang
Huynh, Xuan Thanh
Mai, Phuc Hong
Nguyen, Vinh The
Huynh, Quang Nhut
Nguyen, Huy Tien
Le, Tung
author_facet Nguyen, Minh Hoang
Pham, Vu Hoang
Huynh, Xuan Thanh
Mai, Phuc Hong
Nguyen, Vinh The
Huynh, Quang Nhut
Nguyen, Huy Tien
Le, Tung
contents Large language models (LLMs) have recently reshaped Automated Essay Scoring (AES), yet prior studies typically examine individual techniques in isolation, limiting understanding of their relative merits for English as a Second Language (L2) writing. To bridge this gap, we presents a comprehensive comparison of major LLM-based AES paradigms on IELTS Writing Task~2. On this unified benchmark, we evaluate four approaches: (i) encoder-based classification fine-tuning, (ii) zero- and few-shot prompting, (iii) instruction tuning and Retrieval-Augmented Generation (RAG), and (iv) Supervised Fine-Tuning combined with Direct Preference Optimization (DPO) and RAG. Our results reveal clear accuracy-cost-robustness trade-offs across methods, the best configuration, integrating k-SFT and RAG, achieves the strongest overall results with F1-Score 93%. This study offers the first unified empirical comparison of modern LLM-based AES strategies for English L2, promising potential in auto-grading writing tasks. Code is public at https://github.com/MinhNguyenDS/LLM_AES-EnL2
format Preprint
id arxiv_https___arxiv_org_abs_2603_06424
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring
Nguyen, Minh Hoang
Pham, Vu Hoang
Huynh, Xuan Thanh
Mai, Phuc Hong
Nguyen, Vinh The
Huynh, Quang Nhut
Nguyen, Huy Tien
Le, Tung
Computation and Language
68T50, 68U10, 68T05
I.2.7; I.2.6; K.3.1
Large language models (LLMs) have recently reshaped Automated Essay Scoring (AES), yet prior studies typically examine individual techniques in isolation, limiting understanding of their relative merits for English as a Second Language (L2) writing. To bridge this gap, we presents a comprehensive comparison of major LLM-based AES paradigms on IELTS Writing Task~2. On this unified benchmark, we evaluate four approaches: (i) encoder-based classification fine-tuning, (ii) zero- and few-shot prompting, (iii) instruction tuning and Retrieval-Augmented Generation (RAG), and (iv) Supervised Fine-Tuning combined with Direct Preference Optimization (DPO) and RAG. Our results reveal clear accuracy-cost-robustness trade-offs across methods, the best configuration, integrating k-SFT and RAG, achieves the strongest overall results with F1-Score 93%. This study offers the first unified empirical comparison of modern LLM-based AES strategies for English L2, promising potential in auto-grading writing tasks. Code is public at https://github.com/MinhNguyenDS/LLM_AES-EnL2
title From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring
topic Computation and Language
68T50, 68U10, 68T05
I.2.7; I.2.6; K.3.1
url https://arxiv.org/abs/2603.06424