A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915083062345728 |
|---|---|
| author | Liu, Xiaoqian Du, Yangfan Wang, Jianjin Ge, Yuan Xu, Chen Xiao, Tong Chen, Guocheng Zhu, Jingbo |
| author_facet | Liu, Xiaoqian Du, Yangfan Wang, Jianjin Ge, Yuan Xu, Chen Xiao, Tong Chen, Guocheng Zhu, Jingbo |
| contents | Simultaneous Speech Translation (SimulST) involves generating target language text while continuously processing streaming speech input, presenting significant real-time challenges. Multi-task learning is often employed to enhance SimulST performance but introduces optimization conflicts between primary and auxiliary tasks, potentially compromising overall efficiency. The existing model-level conflict resolution methods are not well-suited for this task which exacerbates inefficiencies and leads to high GPU memory consumption. To address these challenges, we propose a Modular Gradient Conflict Mitigation (MGCM) strategy that detects conflicts at a finer-grained modular level and resolves them utilizing gradient projection. Experimental results demonstrate that MGCM significantly improves SimulST performance, particularly under medium and high latency conditions, achieving a 0.68 BLEU score gain in offline tasks. Additionally, MGCM reduces GPU memory consumption by over 95\% compared to other conflict mitigation methods, establishing it as a robust solution for SimulST tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_15911 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation Liu, Xiaoqian Du, Yangfan Wang, Jianjin Ge, Yuan Xu, Chen Xiao, Tong Chen, Guocheng Zhu, Jingbo Computation and Language Sound Audio and Speech Processing Simultaneous Speech Translation (SimulST) involves generating target language text while continuously processing streaming speech input, presenting significant real-time challenges. Multi-task learning is often employed to enhance SimulST performance but introduces optimization conflicts between primary and auxiliary tasks, potentially compromising overall efficiency. The existing model-level conflict resolution methods are not well-suited for this task which exacerbates inefficiencies and leads to high GPU memory consumption. To address these challenges, we propose a Modular Gradient Conflict Mitigation (MGCM) strategy that detects conflicts at a finer-grained modular level and resolves them utilizing gradient projection. Experimental results demonstrate that MGCM significantly improves SimulST performance, particularly under medium and high latency conditions, achieving a 0.68 BLEU score gain in offline tasks. Additionally, MGCM reduces GPU memory consumption by over 95\% compared to other conflict mitigation methods, establishing it as a robust solution for SimulST tasks. |
| title | A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2409.15911 |