Data Augmentation for Spoken Grammatical Error Correction

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Karanasou, Penny, Qian, Mengjie, Bannò, Stefano, Gales, Mark J. F., Knill, Kate M.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908467455852544
author Karanasou, Penny
Qian, Mengjie
Bannò, Stefano
Gales, Mark J. F.
Knill, Kate M.
author_facet Karanasou, Penny
Qian, Mengjie
Bannò, Stefano
Gales, Mark J. F.
Knill, Kate M.
contents While there exist strong benchmark datasets for grammatical error correction (GEC), high-quality annotated spoken datasets for Spoken GEC (SGEC) are still under-resourced. In this paper, we propose a fully automated method to generate audio-text pairs with grammatical errors and disfluencies. Moreover, we propose a series of objective metrics that can be used to evaluate the generated data and choose the more suitable dataset for SGEC. The goal is to generate an augmented dataset that maintains the textual and acoustic characteristics of the original data while providing new types of errors. This augmented dataset should augment and enrich the original corpus without altering the language assessment scores of the second language (L2) learners. We evaluate the use of the augmented corpus both for written GEC (the text part) and for SGEC (the audio-text pairs). Our experiments are conducted on the S\&I Corpus, the first publicly available speech dataset with grammar error annotations.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19374
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data Augmentation for Spoken Grammatical Error Correction
Karanasou, Penny
Qian, Mengjie
Bannò, Stefano
Gales, Mark J. F.
Knill, Kate M.
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
While there exist strong benchmark datasets for grammatical error correction (GEC), high-quality annotated spoken datasets for Spoken GEC (SGEC) are still under-resourced. In this paper, we propose a fully automated method to generate audio-text pairs with grammatical errors and disfluencies. Moreover, we propose a series of objective metrics that can be used to evaluate the generated data and choose the more suitable dataset for SGEC. The goal is to generate an augmented dataset that maintains the textual and acoustic characteristics of the original data while providing new types of errors. This augmented dataset should augment and enrich the original corpus without altering the language assessment scores of the second language (L2) learners. We evaluate the use of the augmented corpus both for written GEC (the text part) and for SGEC (the audio-text pairs). Our experiments are conducted on the S\&I Corpus, the first publicly available speech dataset with grammar error annotations.
title Data Augmentation for Spoken Grammatical Error Correction
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2507.19374