A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Xiaoqian, Du, Yangfan, Wang, Jianjin, Ge, Yuan, Xu, Chen, Xiao, Tong, Chen, Guocheng, Zhu, Jingbo
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915083062345728
author Liu, Xiaoqian
Du, Yangfan
Wang, Jianjin
Ge, Yuan
Xu, Chen
Xiao, Tong
Chen, Guocheng
Zhu, Jingbo
author_facet Liu, Xiaoqian
Du, Yangfan
Wang, Jianjin
Ge, Yuan
Xu, Chen
Xiao, Tong
Chen, Guocheng
Zhu, Jingbo
contents Simultaneous Speech Translation (SimulST) involves generating target language text while continuously processing streaming speech input, presenting significant real-time challenges. Multi-task learning is often employed to enhance SimulST performance but introduces optimization conflicts between primary and auxiliary tasks, potentially compromising overall efficiency. The existing model-level conflict resolution methods are not well-suited for this task which exacerbates inefficiencies and leads to high GPU memory consumption. To address these challenges, we propose a Modular Gradient Conflict Mitigation (MGCM) strategy that detects conflicts at a finer-grained modular level and resolves them utilizing gradient projection. Experimental results demonstrate that MGCM significantly improves SimulST performance, particularly under medium and high latency conditions, achieving a 0.68 BLEU score gain in offline tasks. Additionally, MGCM reduces GPU memory consumption by over 95\% compared to other conflict mitigation methods, establishing it as a robust solution for SimulST tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2409_15911
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation
Liu, Xiaoqian
Du, Yangfan
Wang, Jianjin
Ge, Yuan
Xu, Chen
Xiao, Tong
Chen, Guocheng
Zhu, Jingbo
Computation and Language
Sound
Audio and Speech Processing
Simultaneous Speech Translation (SimulST) involves generating target language text while continuously processing streaming speech input, presenting significant real-time challenges. Multi-task learning is often employed to enhance SimulST performance but introduces optimization conflicts between primary and auxiliary tasks, potentially compromising overall efficiency. The existing model-level conflict resolution methods are not well-suited for this task which exacerbates inefficiencies and leads to high GPU memory consumption. To address these challenges, we propose a Modular Gradient Conflict Mitigation (MGCM) strategy that detects conflicts at a finer-grained modular level and resolves them utilizing gradient projection. Experimental results demonstrate that MGCM significantly improves SimulST performance, particularly under medium and high latency conditions, achieving a 0.68 BLEU score gain in offline tasks. Additionally, MGCM reduces GPU memory consumption by over 95\% compared to other conflict mitigation methods, establishing it as a robust solution for SimulST tasks.
title A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2409.15911