TICL+: A Case Study On Speech In-Context Learning for Children's Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Haolong, Yegorova, Yekaterina, Hasegawa-Johnson, Mark
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909972344864768
author Zheng, Haolong
Yegorova, Yekaterina
Hasegawa-Johnson, Mark
author_facet Zheng, Haolong
Yegorova, Yekaterina
Hasegawa-Johnson, Mark
contents Children's speech recognition remains challenging due to substantial acoustic and linguistic variability, limited labeled data, and significant differences from adult speech. Speech foundation models can address these challenges through Speech In-Context Learning (SICL), allowing adaptation to new domains without fine-tuning. However, the effectiveness of SICL depends on how in-context examples are selected. We extend an existing retrieval-based method, Text-Embedding KNN for SICL (TICL), introducing an acoustic reranking step to create TICL+. This extension prioritizes examples that are both semantically and acoustically aligned with the test input. Experiments on four children's speech corpora show that TICL+ achieves up to a 53.3% relative word error rate reduction over zero-shot performance and 37.6% over baseline TICL, highlighting the value of combining semantic and acoustic information for robust, scalable ASR in children's speech.
format Preprint
id arxiv_https___arxiv_org_abs_2512_18263
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TICL+: A Case Study On Speech In-Context Learning for Children's Speech Recognition
Zheng, Haolong
Yegorova, Yekaterina
Hasegawa-Johnson, Mark
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
Children's speech recognition remains challenging due to substantial acoustic and linguistic variability, limited labeled data, and significant differences from adult speech. Speech foundation models can address these challenges through Speech In-Context Learning (SICL), allowing adaptation to new domains without fine-tuning. However, the effectiveness of SICL depends on how in-context examples are selected. We extend an existing retrieval-based method, Text-Embedding KNN for SICL (TICL), introducing an acoustic reranking step to create TICL+. This extension prioritizes examples that are both semantically and acoustically aligned with the test input. Experiments on four children's speech corpora show that TICL+ achieves up to a 53.3% relative word error rate reduction over zero-shot performance and 37.6% over baseline TICL, highlighting the value of combining semantic and acoustic information for robust, scalable ASR in children's speech.
title TICL+: A Case Study On Speech In-Context Learning for Children's Speech Recognition
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2512.18263