Continual Speech Learning with Fused Speech Features
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910982151864320 |
|---|---|
| author | Wang, Guitao Zhao, Jinming Yang, Hao Qi, Guilin Wu, Tongtong Haffari, Gholamreza |
| author_facet | Wang, Guitao Zhao, Jinming Yang, Hao Qi, Guilin Wu, Tongtong Haffari, Gholamreza |
| contents | Rapid growth in speech data demands adaptive models, as traditional static methods fail to keep pace with dynamic and diverse speech information. We introduce continuous speech learning, a new set-up targeting at bridging the adaptation gap in current speech models. We use the encoder-decoder Whisper model to standardize speech tasks into a generative format. We integrate a learnable gated-fusion layer on the top of the encoder to dynamically select task-specific features for downstream tasks. Our approach improves accuracy significantly over traditional methods in six speech processing tasks, demonstrating gains in adapting to new speech tasks without full retraining. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_01496 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Continual Speech Learning with Fused Speech Features Wang, Guitao Zhao, Jinming Yang, Hao Qi, Guilin Wu, Tongtong Haffari, Gholamreza Computation and Language Sound Audio and Speech Processing Rapid growth in speech data demands adaptive models, as traditional static methods fail to keep pace with dynamic and diverse speech information. We introduce continuous speech learning, a new set-up targeting at bridging the adaptation gap in current speech models. We use the encoder-decoder Whisper model to standardize speech tasks into a generative format. We integrate a learnable gated-fusion layer on the top of the encoder to dynamically select task-specific features for downstream tasks. Our approach improves accuracy significantly over traditional methods in six speech processing tasks, demonstrating gains in adapting to new speech tasks without full retraining. |
| title | Continual Speech Learning with Fused Speech Features |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.01496 |