Continual Learning in Machine Speech Chain Using Gradient Episodic Memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tyndall, Geoffrey, Azizah, Kurniawati, Tanaya, Dipta, Purwarianti, Ayu, Lestari, Dessi Puji, Sakti, Sakriani
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910718722310144
author Tyndall, Geoffrey
Azizah, Kurniawati
Tanaya, Dipta
Purwarianti, Ayu
Lestari, Dessi Puji
Sakti, Sakriani
author_facet Tyndall, Geoffrey
Azizah, Kurniawati
Tanaya, Dipta
Purwarianti, Ayu
Lestari, Dessi Puji
Sakti, Sakriani
contents Continual learning for automatic speech recognition (ASR) systems poses a challenge, especially with the need to avoid catastrophic forgetting while maintaining performance on previously learned tasks. This paper introduces a novel approach leveraging the machine speech chain framework to enable continual learning in ASR using gradient episodic memory (GEM). By incorporating a text-to-speech (TTS) component within the machine speech chain, we support the replay mechanism essential for GEM, allowing the ASR model to learn new tasks sequentially without significant performance degradation on earlier tasks. Our experiments, conducted on the LJ Speech dataset, demonstrate that our method outperforms traditional fine-tuning and multitask learning approaches, achieving a substantial error rate reduction while maintaining high performance across varying noise conditions. We showed the potential of our semi-supervised machine speech chain approach for effective and efficient continual learning in speech recognition.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18320
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Continual Learning in Machine Speech Chain Using Gradient Episodic Memory
Tyndall, Geoffrey
Azizah, Kurniawati
Tanaya, Dipta
Purwarianti, Ayu
Lestari, Dessi Puji
Sakti, Sakriani
Computation and Language
Artificial Intelligence
Audio and Speech Processing
Continual learning for automatic speech recognition (ASR) systems poses a challenge, especially with the need to avoid catastrophic forgetting while maintaining performance on previously learned tasks. This paper introduces a novel approach leveraging the machine speech chain framework to enable continual learning in ASR using gradient episodic memory (GEM). By incorporating a text-to-speech (TTS) component within the machine speech chain, we support the replay mechanism essential for GEM, allowing the ASR model to learn new tasks sequentially without significant performance degradation on earlier tasks. Our experiments, conducted on the LJ Speech dataset, demonstrate that our method outperforms traditional fine-tuning and multitask learning approaches, achieving a substantial error rate reduction while maintaining high performance across varying noise conditions. We showed the potential of our semi-supervised machine speech chain approach for effective and efficient continual learning in speech recognition.
title Continual Learning in Machine Speech Chain Using Gradient Episodic Memory
topic Computation and Language
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2411.18320