Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hsiao, Chi-Yuan, Lu, Ke-Han, Chang, Kai-Wei, Yang, Chih-Kai, Chen, Wei-Chih, Lee, Hung-yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915300616699904
author Hsiao, Chi-Yuan
Lu, Ke-Han
Chang, Kai-Wei
Yang, Chih-Kai
Chen, Wei-Chih
Lee, Hung-yi
author_facet Hsiao, Chi-Yuan
Lu, Ke-Han
Chang, Kai-Wei
Yang, Chih-Kai
Chen, Wei-Chih
Lee, Hung-yi
contents End-to-end training of Spoken Language Models (SLMs) commonly involves adapting pre-trained text-based Large Language Models (LLMs) to the speech modality through multi-stage training on diverse tasks such as ASR, TTS and spoken question answering (SQA). Although this multi-stage continual learning equips LLMs with both speech understanding and generation capabilities, the substantial differences in task and data distributions across stages can lead to catastrophic forgetting, where previously acquired knowledge is lost. This paper investigates catastrophic forgetting and evaluates three mitigation strategies-model merging, discounting the LoRA scaling factor, and experience replay to balance knowledge retention with new learning. Results show that experience replay is the most effective, with further gains achieved by combining it with other methods. These findings provide insights for developing more robust and efficient SLM training pipelines.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17496
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
Hsiao, Chi-Yuan
Lu, Ke-Han
Chang, Kai-Wei
Yang, Chih-Kai
Chen, Wei-Chih
Lee, Hung-yi
Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
End-to-end training of Spoken Language Models (SLMs) commonly involves adapting pre-trained text-based Large Language Models (LLMs) to the speech modality through multi-stage training on diverse tasks such as ASR, TTS and spoken question answering (SQA). Although this multi-stage continual learning equips LLMs with both speech understanding and generation capabilities, the substantial differences in task and data distributions across stages can lead to catastrophic forgetting, where previously acquired knowledge is lost. This paper investigates catastrophic forgetting and evaluates three mitigation strategies-model merging, discounting the LoRA scaling factor, and experience replay to balance knowledge retention with new learning. Results show that experience replay is the most effective, with further gains achieved by combining it with other methods. These findings provide insights for developing more robust and efficient SLM training pipelines.
title Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.17496