Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915300616699904 |
|---|---|
| author | Hsiao, Chi-Yuan Lu, Ke-Han Chang, Kai-Wei Yang, Chih-Kai Chen, Wei-Chih Lee, Hung-yi |
| author_facet | Hsiao, Chi-Yuan Lu, Ke-Han Chang, Kai-Wei Yang, Chih-Kai Chen, Wei-Chih Lee, Hung-yi |
| contents | End-to-end training of Spoken Language Models (SLMs) commonly involves adapting pre-trained text-based Large Language Models (LLMs) to the speech modality through multi-stage training on diverse tasks such as ASR, TTS and spoken question answering (SQA). Although this multi-stage continual learning equips LLMs with both speech understanding and generation capabilities, the substantial differences in task and data distributions across stages can lead to catastrophic forgetting, where previously acquired knowledge is lost. This paper investigates catastrophic forgetting and evaluates three mitigation strategies-model merging, discounting the LoRA scaling factor, and experience replay to balance knowledge retention with new learning. Results show that experience replay is the most effective, with further gains achieved by combining it with other methods. These findings provide insights for developing more robust and efficient SLM training pipelines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_17496 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models Hsiao, Chi-Yuan Lu, Ke-Han Chang, Kai-Wei Yang, Chih-Kai Chen, Wei-Chih Lee, Hung-yi Computation and Language Artificial Intelligence Machine Learning Sound Audio and Speech Processing End-to-end training of Spoken Language Models (SLMs) commonly involves adapting pre-trained text-based Large Language Models (LLMs) to the speech modality through multi-stage training on diverse tasks such as ASR, TTS and spoken question answering (SQA). Although this multi-stage continual learning equips LLMs with both speech understanding and generation capabilities, the substantial differences in task and data distributions across stages can lead to catastrophic forgetting, where previously acquired knowledge is lost. This paper investigates catastrophic forgetting and evaluates three mitigation strategies-model merging, discounting the LoRA scaling factor, and experience replay to balance knowledge retention with new learning. Results show that experience replay is the most effective, with further gains achieved by combining it with other methods. These findings provide insights for developing more robust and efficient SLM training pipelines. |
| title | Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models |
| topic | Computation and Language Artificial Intelligence Machine Learning Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2505.17496 |