Natural Language-based Assessment of L2 Oral Proficiency using LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bannò, Stefano, Ma, Rao, Qian, Mengjie, Tang, Siyuan, Knill, Kate, Gales, Mark
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915388762095616
author Bannò, Stefano
Ma, Rao
Qian, Mengjie
Tang, Siyuan
Knill, Kate
Gales, Mark
author_facet Bannò, Stefano
Ma, Rao
Qian, Mengjie
Tang, Siyuan
Knill, Kate
Gales, Mark
contents Natural language-based assessment (NLA) is an approach to second language assessment that uses instructions - expressed in the form of can-do descriptors - originally intended for human examiners, aiming to determine whether large language models (LLMs) can interpret and apply them in ways comparable to human assessment. In this work, we explore the use of such descriptors with an open-source LLM, Qwen 2.5 72B, to assess responses from the publicly available S&I Corpus in a zero-shot setting. Our results show that this approach - relying solely on textual information - achieves competitive performance: while it does not outperform state-of-the-art speech LLMs fine-tuned for the task, it surpasses a BERT-based model trained specifically for this purpose. NLA proves particularly effective in mismatched task settings, is generalisable to other data types and languages, and offers greater interpretability, as it is grounded in clearly explainable, widely applicable language descriptors.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10200
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Natural Language-based Assessment of L2 Oral Proficiency using LLMs
Bannò, Stefano
Ma, Rao
Qian, Mengjie
Tang, Siyuan
Knill, Kate
Gales, Mark
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Natural language-based assessment (NLA) is an approach to second language assessment that uses instructions - expressed in the form of can-do descriptors - originally intended for human examiners, aiming to determine whether large language models (LLMs) can interpret and apply them in ways comparable to human assessment. In this work, we explore the use of such descriptors with an open-source LLM, Qwen 2.5 72B, to assess responses from the publicly available S&I Corpus in a zero-shot setting. Our results show that this approach - relying solely on textual information - achieves competitive performance: while it does not outperform state-of-the-art speech LLMs fine-tuned for the task, it surpasses a BERT-based model trained specifically for this purpose. NLA proves particularly effective in mismatched task settings, is generalisable to other data types and languages, and offers greater interpretability, as it is grounded in clearly explainable, widely applicable language descriptors.
title Natural Language-based Assessment of L2 Oral Proficiency using LLMs
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2507.10200