Dynamic Data Pruning for Automatic Speech Recognition

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xiao, Qiao, Ma, Pingchuan, Fernandez-Lopez, Adriana, Wu, Boqian, Yin, Lu, Petridis, Stavros, Pechenizkiy, Mykola, Pantic, Maja, Mocanu, Decebal Constantin, Liu, Shiwei
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866907958223306752
author Xiao, Qiao
Ma, Pingchuan
Fernandez-Lopez, Adriana
Wu, Boqian
Yin, Lu
Petridis, Stavros
Pechenizkiy, Mykola
Pantic, Maja
Mocanu, Decebal Constantin
Liu, Shiwei
author_facet Xiao, Qiao
Ma, Pingchuan
Fernandez-Lopez, Adriana
Wu, Boqian
Yin, Lu
Petridis, Stavros
Pechenizkiy, Mykola
Pantic, Maja
Mocanu, Decebal Constantin
Liu, Shiwei
contents The recent success of Automatic Speech Recognition (ASR) is largely attributed to the ever-growing amount of training data. However, this trend has made model training prohibitively costly and imposed computational demands. While data pruning has been proposed to mitigate this issue by identifying a small subset of relevant data, its application in ASR has been barely explored, and existing works often entail significant overhead to achieve meaningful results. To fill this gap, this paper presents the first investigation of dynamic data pruning for ASR, finding that we can reach the full-data performance by dynamically selecting 70% of data. Furthermore, we introduce Dynamic Data Pruning for ASR (DDP-ASR), which offers several fine-grained pruning granularities specifically tailored for speech-related datasets, going beyond the conventional pruning of entire time sequences. Our intensive experiments show that DDP-ASR can save up to 1.6x training time with negligible performance loss.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18373
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dynamic Data Pruning for Automatic Speech Recognition
Xiao, Qiao
Ma, Pingchuan
Fernandez-Lopez, Adriana
Wu, Boqian
Yin, Lu
Petridis, Stavros
Pechenizkiy, Mykola
Pantic, Maja
Mocanu, Decebal Constantin
Liu, Shiwei
Computation and Language
Sound
Audio and Speech Processing
The recent success of Automatic Speech Recognition (ASR) is largely attributed to the ever-growing amount of training data. However, this trend has made model training prohibitively costly and imposed computational demands. While data pruning has been proposed to mitigate this issue by identifying a small subset of relevant data, its application in ASR has been barely explored, and existing works often entail significant overhead to achieve meaningful results. To fill this gap, this paper presents the first investigation of dynamic data pruning for ASR, finding that we can reach the full-data performance by dynamically selecting 70% of data. Furthermore, we introduce Dynamic Data Pruning for ASR (DDP-ASR), which offers several fine-grained pruning granularities specifically tailored for speech-related datasets, going beyond the conventional pruning of entire time sequences. Our intensive experiments show that DDP-ASR can save up to 1.6x training time with negligible performance loss.
title Dynamic Data Pruning for Automatic Speech Recognition
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.18373