Probing the Hidden Talent of ASR Foundation Models for L2 English Oral Assessment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chao, Fu-An, Yan, Bi-Cheng, Chen, Berlin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910000545267712
author Chao, Fu-An
Yan, Bi-Cheng
Chen, Berlin
author_facet Chao, Fu-An
Yan, Bi-Cheng
Chen, Berlin
contents In this paper, we explore the untapped potential of Whisper, a well-established automatic speech recognition (ASR) foundation model, in the context of L2 spoken language assessment (SLA). Unlike prior studies that extrinsically analyze transcriptions produced by Whisper, our approach goes a step further to probe its latent capabilities by extracting acoustic and linguistic features from hidden representations. With only a lightweight classifier being trained on top of Whisper's intermediate and final outputs, our method achieves strong performance on the GEPT picture-description dataset, outperforming existing cutting-edge baselines, including a multimodal approach. Furthermore, by incorporating image and text-prompt information as auxiliary relevance cues, we demonstrate additional performance gains. Finally, we conduct an in-depth analysis of Whisper's embeddings, which reveals that, even without task-specific fine-tuning, the model intrinsically encodes both ordinal proficiency patterns and semantic aspects of speech, highlighting its potential as a powerful foundation for SLA and other spoken language understanding tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16387
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Probing the Hidden Talent of ASR Foundation Models for L2 English Oral Assessment
Chao, Fu-An
Yan, Bi-Cheng
Chen, Berlin
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
In this paper, we explore the untapped potential of Whisper, a well-established automatic speech recognition (ASR) foundation model, in the context of L2 spoken language assessment (SLA). Unlike prior studies that extrinsically analyze transcriptions produced by Whisper, our approach goes a step further to probe its latent capabilities by extracting acoustic and linguistic features from hidden representations. With only a lightweight classifier being trained on top of Whisper's intermediate and final outputs, our method achieves strong performance on the GEPT picture-description dataset, outperforming existing cutting-edge baselines, including a multimodal approach. Furthermore, by incorporating image and text-prompt information as auxiliary relevance cues, we demonstrate additional performance gains. Finally, we conduct an in-depth analysis of Whisper's embeddings, which reveals that, even without task-specific fine-tuning, the model intrinsically encodes both ordinal proficiency patterns and semantic aspects of speech, highlighting its potential as a powerful foundation for SLA and other spoken language understanding tasks.
title Probing the Hidden Talent of ASR Foundation Models for L2 English Oral Assessment
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2510.16387