Optimizing Speech Multi-View Feature Fusion through Conditional Computation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866912188347711488 |
|---|---|
| author | Shan, Weiqiao Zhang, Yuhao Han, Yuchen Li, Bei Zhao, Xiaofeng Li, Yuang Zhang, Min Yang, Hao Xiao, Tong Zhu, Jingbo |
| author_facet | Shan, Weiqiao Zhang, Yuhao Han, Yuchen Li, Bei Zhao, Xiaofeng Li, Yuang Zhang, Min Yang, Hao Xiao, Tong Zhu, Jingbo |
| contents | Recent advancements have highlighted the efficacy of self-supervised learning (SSL) features in various speech-related tasks, providing lightweight and versatile multi-view speech representations. However, our study reveals that while SSL features expedite model convergence, they conflict with traditional spectral features like FBanks in terms of update directions. In response, we propose a novel generalized feature fusion framework grounded in conditional computation, featuring a gradient-sensitive gating network and a multi-stage dropout strategy. This framework mitigates feature conflicts and bolsters model robustness to multi-view input features. By integrating SSL and spectral features, our approach accelerates convergence and maintains performance on par with spectral models across multiple speech translation tasks on the MUSTC dataset. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_08057 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Optimizing Speech Multi-View Feature Fusion through Conditional Computation Shan, Weiqiao Zhang, Yuhao Han, Yuchen Li, Bei Zhao, Xiaofeng Li, Yuang Zhang, Min Yang, Hao Xiao, Tong Zhu, Jingbo Audio and Speech Processing Artificial Intelligence Computation and Language Sound Recent advancements have highlighted the efficacy of self-supervised learning (SSL) features in various speech-related tasks, providing lightweight and versatile multi-view speech representations. However, our study reveals that while SSL features expedite model convergence, they conflict with traditional spectral features like FBanks in terms of update directions. In response, we propose a novel generalized feature fusion framework grounded in conditional computation, featuring a gradient-sensitive gating network and a multi-stage dropout strategy. This framework mitigates feature conflicts and bolsters model robustness to multi-view input features. By integrating SSL and spectral features, our approach accelerates convergence and maintains performance on par with spectral models across multiple speech translation tasks on the MUSTC dataset. |
| title | Optimizing Speech Multi-View Feature Fusion through Conditional Computation |
| topic | Audio and Speech Processing Artificial Intelligence Computation and Language Sound |
| url | https://arxiv.org/abs/2501.08057 |