Large Language Models based ASR Error Correction for Child Conversations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Anfeng, Feng, Tiantian, Kim, So Hyun, Bishop, Somer, Lord, Catherine, Narayanan, Shrikanth
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915303117553664
author Xu, Anfeng
Feng, Tiantian
Kim, So Hyun
Bishop, Somer
Lord, Catherine
Narayanan, Shrikanth
author_facet Xu, Anfeng
Feng, Tiantian
Kim, So Hyun
Bishop, Somer
Lord, Catherine
Narayanan, Shrikanth
contents Automatic Speech Recognition (ASR) has recently shown remarkable progress, but accurately transcribing children's speech remains a significant challenge. Recent developments in Large Language Models (LLMs) have shown promise in improving ASR transcriptions. However, their applications in child speech including conversational scenarios are underexplored. In this study, we explore the use of LLMs in correcting ASR errors for conversational child speech. We demonstrate the promises and challenges of LLMs through experiments on two children's conversational speech datasets with both zero-shot and fine-tuned ASR outputs. We find that while LLMs are helpful in correcting zero-shot ASR outputs and fine-tuned CTC-based ASR outputs, it remains challenging for LLMs to improve ASR performance when incorporating contextual information or when using fine-tuned autoregressive ASR (e.g., Whisper) outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16212
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Models based ASR Error Correction for Child Conversations
Xu, Anfeng
Feng, Tiantian
Kim, So Hyun
Bishop, Somer
Lord, Catherine
Narayanan, Shrikanth
Computation and Language
Audio and Speech Processing
Automatic Speech Recognition (ASR) has recently shown remarkable progress, but accurately transcribing children's speech remains a significant challenge. Recent developments in Large Language Models (LLMs) have shown promise in improving ASR transcriptions. However, their applications in child speech including conversational scenarios are underexplored. In this study, we explore the use of LLMs in correcting ASR errors for conversational child speech. We demonstrate the promises and challenges of LLMs through experiments on two children's conversational speech datasets with both zero-shot and fine-tuned ASR outputs. We find that while LLMs are helpful in correcting zero-shot ASR outputs and fine-tuned CTC-based ASR outputs, it remains challenging for LLMs to improve ASR performance when incorporating contextual information or when using fine-tuned autoregressive ASR (e.g., Whisper) outputs.
title Large Language Models based ASR Error Correction for Child Conversations
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2505.16212