Optimizing Speech Multi-View Feature Fusion through Conditional Computation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shan, Weiqiao, Zhang, Yuhao, Han, Yuchen, Li, Bei, Zhao, Xiaofeng, Li, Yuang, Zhang, Min, Yang, Hao, Xiao, Tong, Zhu, Jingbo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912188347711488
author Shan, Weiqiao
Zhang, Yuhao
Han, Yuchen
Li, Bei
Zhao, Xiaofeng
Li, Yuang
Zhang, Min
Yang, Hao
Xiao, Tong
Zhu, Jingbo
author_facet Shan, Weiqiao
Zhang, Yuhao
Han, Yuchen
Li, Bei
Zhao, Xiaofeng
Li, Yuang
Zhang, Min
Yang, Hao
Xiao, Tong
Zhu, Jingbo
contents Recent advancements have highlighted the efficacy of self-supervised learning (SSL) features in various speech-related tasks, providing lightweight and versatile multi-view speech representations. However, our study reveals that while SSL features expedite model convergence, they conflict with traditional spectral features like FBanks in terms of update directions. In response, we propose a novel generalized feature fusion framework grounded in conditional computation, featuring a gradient-sensitive gating network and a multi-stage dropout strategy. This framework mitigates feature conflicts and bolsters model robustness to multi-view input features. By integrating SSL and spectral features, our approach accelerates convergence and maintains performance on par with spectral models across multiple speech translation tasks on the MUSTC dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2501_08057
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimizing Speech Multi-View Feature Fusion through Conditional Computation
Shan, Weiqiao
Zhang, Yuhao
Han, Yuchen
Li, Bei
Zhao, Xiaofeng
Li, Yuang
Zhang, Min
Yang, Hao
Xiao, Tong
Zhu, Jingbo
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
Recent advancements have highlighted the efficacy of self-supervised learning (SSL) features in various speech-related tasks, providing lightweight and versatile multi-view speech representations. However, our study reveals that while SSL features expedite model convergence, they conflict with traditional spectral features like FBanks in terms of update directions. In response, we propose a novel generalized feature fusion framework grounded in conditional computation, featuring a gradient-sensitive gating network and a multi-stage dropout strategy. This framework mitigates feature conflicts and bolsters model robustness to multi-view input features. By integrating SSL and spectral features, our approach accelerates convergence and maintains performance on par with spectral models across multiple speech translation tasks on the MUSTC dataset.
title Optimizing Speech Multi-View Feature Fusion through Conditional Computation
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Sound
url https://arxiv.org/abs/2501.08057