Continual Speech Learning with Fused Speech Features

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Guitao, Zhao, Jinming, Yang, Hao, Qi, Guilin, Wu, Tongtong, Haffari, Gholamreza
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910982151864320
author Wang, Guitao
Zhao, Jinming
Yang, Hao
Qi, Guilin
Wu, Tongtong
Haffari, Gholamreza
author_facet Wang, Guitao
Zhao, Jinming
Yang, Hao
Qi, Guilin
Wu, Tongtong
Haffari, Gholamreza
contents Rapid growth in speech data demands adaptive models, as traditional static methods fail to keep pace with dynamic and diverse speech information. We introduce continuous speech learning, a new set-up targeting at bridging the adaptation gap in current speech models. We use the encoder-decoder Whisper model to standardize speech tasks into a generative format. We integrate a learnable gated-fusion layer on the top of the encoder to dynamically select task-specific features for downstream tasks. Our approach improves accuracy significantly over traditional methods in six speech processing tasks, demonstrating gains in adapting to new speech tasks without full retraining.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01496
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Continual Speech Learning with Fused Speech Features
Wang, Guitao
Zhao, Jinming
Yang, Hao
Qi, Guilin
Wu, Tongtong
Haffari, Gholamreza
Computation and Language
Sound
Audio and Speech Processing
Rapid growth in speech data demands adaptive models, as traditional static methods fail to keep pace with dynamic and diverse speech information. We introduce continuous speech learning, a new set-up targeting at bridging the adaptation gap in current speech models. We use the encoder-decoder Whisper model to standardize speech tasks into a generative format. We integrate a learnable gated-fusion layer on the top of the encoder to dynamically select task-specific features for downstream tasks. Our approach improves accuracy significantly over traditional methods in six speech processing tasks, demonstrating gains in adapting to new speech tasks without full retraining.
title Continual Speech Learning with Fused Speech Features
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.01496