Length Aware Speech Translation for Video Dubbing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916771670261760 |
|---|---|
| author | Chadha, Harveen Singh Subramanian, Aswin Shanmugam Joshi, Vikas Bansal, Shubham Xue, Jian Mehta, Rupeshkumar Li, Jinyu |
| author_facet | Chadha, Harveen Singh Subramanian, Aswin Shanmugam Joshi, Vikas Bansal, Shubham Xue, Jian Mehta, Rupeshkumar Li, Jinyu |
| contents | In video dubbing, aligning translated audio with the source audio is a significant challenge. Our focus is on achieving this efficiently, tailored for real-time, on-device video dubbing scenarios. We developed a phoneme-based end-to-end length-sensitive speech translation (LSST) model, which generates translations of varying lengths short, normal, and long using predefined tags. Additionally, we introduced length-aware beam search (LABS), an efficient approach to generate translations of different lengths in a single decoding pass. This approach maintained comparable BLEU scores compared to a baseline without length awareness while significantly enhancing synchronization quality between source and target audio, achieving a mean opinion score (MOS) gain of 0.34 for Spanish and 0.65 for Korean, respectively. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_00740 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Length Aware Speech Translation for Video Dubbing Chadha, Harveen Singh Subramanian, Aswin Shanmugam Joshi, Vikas Bansal, Shubham Xue, Jian Mehta, Rupeshkumar Li, Jinyu Computation and Language Artificial Intelligence Sound Audio and Speech Processing In video dubbing, aligning translated audio with the source audio is a significant challenge. Our focus is on achieving this efficiently, tailored for real-time, on-device video dubbing scenarios. We developed a phoneme-based end-to-end length-sensitive speech translation (LSST) model, which generates translations of varying lengths short, normal, and long using predefined tags. Additionally, we introduced length-aware beam search (LABS), an efficient approach to generate translations of different lengths in a single decoding pass. This approach maintained comparable BLEU scores compared to a baseline without length awareness while significantly enhancing synchronization quality between source and target audio, achieving a mean opinion score (MOS) gain of 0.34 for Spanish and 0.65 for Korean, respectively. |
| title | Length Aware Speech Translation for Video Dubbing |
| topic | Computation and Language Artificial Intelligence Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.00740 |