Which Data Matter? Embedding-Based Data Selection for Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aldeneh, Zakaria, Seto, Skyler, de Seyssel, Maureen, Chi, Jie, Gu, Zijin, Higuchi, Takuya, Jung, Jee-weon, Watanabe, Shinji, Grangier, David, Theobald, Barry-John, Likhomanenko, Tatiana
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910051357163520
author Aldeneh, Zakaria
Seto, Skyler
de Seyssel, Maureen
Chi, Jie
Gu, Zijin
Higuchi, Takuya
Jung, Jee-weon
Watanabe, Shinji
Grangier, David
Theobald, Barry-John
Likhomanenko, Tatiana
author_facet Aldeneh, Zakaria
Seto, Skyler
de Seyssel, Maureen
Chi, Jie
Gu, Zijin
Higuchi, Takuya
Jung, Jee-weon
Watanabe, Shinji
Grangier, David
Theobald, Barry-John
Likhomanenko, Tatiana
contents Modern ASR systems are typically trained on large-scale pseudo-labeled, in-the-wild data spanning multiple domains. While such heterogeneous data benefit generalist models designed for broad deployment, they pose challenges for specialist models targeting specific domains: specialist models lack the capacity to learn from all available data, and one must pay closer attention to addressing the mismatch between training and test conditions. In this work, we study targeted data selection as a strategy to address these challenges, selecting relevant subsets from 100k hours of in-the-wild training data to optimize performance on target domains. We represent speech samples using embeddings that capture complementary characteristic--speaker attributes, phonetic content, and semantic meaning--and analyze how relevance and diversity along these axes when performing data selection affect downstream ASR performance. Our experiments with CTC-based Conformer models show that training on a strategically selected 5% subset can exceed the performance of models trained on the full dataset by up to 36.8% relative WER reduction on target domains.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05819
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Which Data Matter? Embedding-Based Data Selection for Speech Recognition
Aldeneh, Zakaria
Seto, Skyler
de Seyssel, Maureen
Chi, Jie
Gu, Zijin
Higuchi, Takuya
Jung, Jee-weon
Watanabe, Shinji
Grangier, David
Theobald, Barry-John
Likhomanenko, Tatiana
Sound
Modern ASR systems are typically trained on large-scale pseudo-labeled, in-the-wild data spanning multiple domains. While such heterogeneous data benefit generalist models designed for broad deployment, they pose challenges for specialist models targeting specific domains: specialist models lack the capacity to learn from all available data, and one must pay closer attention to addressing the mismatch between training and test conditions. In this work, we study targeted data selection as a strategy to address these challenges, selecting relevant subsets from 100k hours of in-the-wild training data to optimize performance on target domains. We represent speech samples using embeddings that capture complementary characteristic--speaker attributes, phonetic content, and semantic meaning--and analyze how relevance and diversity along these axes when performing data selection affect downstream ASR performance. Our experiments with CTC-based Conformer models show that training on a strategically selected 5% subset can exceed the performance of models trained on the full dataset by up to 36.8% relative WER reduction on target domains.
title Which Data Matter? Embedding-Based Data Selection for Speech Recognition
topic Sound
url https://arxiv.org/abs/2603.05819