Towards End-to-End Spoken Grammatical Error Correction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bannò, Stefano, Ma, Rao, Qian, Mengjie, Knill, Kate M., Gales, Mark J. F.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909261920993280
author Bannò, Stefano
Ma, Rao
Qian, Mengjie
Knill, Kate M.
Gales, Mark J. F.
author_facet Bannò, Stefano
Ma, Rao
Qian, Mengjie
Knill, Kate M.
Gales, Mark J. F.
contents Grammatical feedback is crucial for L2 learners, teachers, and testers. Spoken grammatical error correction (GEC) aims to supply feedback to L2 learners on their use of grammar when speaking. This process usually relies on a cascaded pipeline comprising an ASR system, disfluency removal, and GEC, with the associated concern of propagating errors between these individual modules. In this paper, we introduce an alternative "end-to-end" approach to spoken GEC, exploiting a speech recognition foundation model, Whisper. This foundation model can be used to replace the whole framework or part of it, e.g., ASR and disfluency removal. These end-to-end approaches are compared to more standard cascaded approaches on the data obtained from a free-speaking spoken language assessment test, Linguaskill. Results demonstrate that end-to-end spoken GEC is possible within this architecture, but the lack of available data limits current performance compared to a system using large quantities of text-based GEC data. Conversely, end-to-end disfluency detection and removal, which is easier for the attention-based Whisper to learn, does outperform cascaded approaches. Additionally, the paper discusses the challenges of providing feedback to candidates when using end-to-end systems for spoken GEC.
format Preprint
id arxiv_https___arxiv_org_abs_2311_05550
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Towards End-to-End Spoken Grammatical Error Correction
Bannò, Stefano
Ma, Rao
Qian, Mengjie
Knill, Kate M.
Gales, Mark J. F.
Computation and Language
Machine Learning
Audio and Speech Processing
Grammatical feedback is crucial for L2 learners, teachers, and testers. Spoken grammatical error correction (GEC) aims to supply feedback to L2 learners on their use of grammar when speaking. This process usually relies on a cascaded pipeline comprising an ASR system, disfluency removal, and GEC, with the associated concern of propagating errors between these individual modules. In this paper, we introduce an alternative "end-to-end" approach to spoken GEC, exploiting a speech recognition foundation model, Whisper. This foundation model can be used to replace the whole framework or part of it, e.g., ASR and disfluency removal. These end-to-end approaches are compared to more standard cascaded approaches on the data obtained from a free-speaking spoken language assessment test, Linguaskill. Results demonstrate that end-to-end spoken GEC is possible within this architecture, but the lack of available data limits current performance compared to a system using large quantities of text-based GEC data. Conversely, end-to-end disfluency detection and removal, which is easier for the attention-based Whisper to learn, does outperform cascaded approaches. Additionally, the paper discusses the challenges of providing feedback to candidates when using end-to-end systems for spoken GEC.
title Towards End-to-End Spoken Grammatical Error Correction
topic Computation and Language
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2311.05550