Length Aware Speech Translation for Video Dubbing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chadha, Harveen Singh, Subramanian, Aswin Shanmugam, Joshi, Vikas, Bansal, Shubham, Xue, Jian, Mehta, Rupeshkumar, Li, Jinyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916771670261760
author Chadha, Harveen Singh
Subramanian, Aswin Shanmugam
Joshi, Vikas
Bansal, Shubham
Xue, Jian
Mehta, Rupeshkumar
Li, Jinyu
author_facet Chadha, Harveen Singh
Subramanian, Aswin Shanmugam
Joshi, Vikas
Bansal, Shubham
Xue, Jian
Mehta, Rupeshkumar
Li, Jinyu
contents In video dubbing, aligning translated audio with the source audio is a significant challenge. Our focus is on achieving this efficiently, tailored for real-time, on-device video dubbing scenarios. We developed a phoneme-based end-to-end length-sensitive speech translation (LSST) model, which generates translations of varying lengths short, normal, and long using predefined tags. Additionally, we introduced length-aware beam search (LABS), an efficient approach to generate translations of different lengths in a single decoding pass. This approach maintained comparable BLEU scores compared to a baseline without length awareness while significantly enhancing synchronization quality between source and target audio, achieving a mean opinion score (MOS) gain of 0.34 for Spanish and 0.65 for Korean, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00740
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Length Aware Speech Translation for Video Dubbing
Chadha, Harveen Singh
Subramanian, Aswin Shanmugam
Joshi, Vikas
Bansal, Shubham
Xue, Jian
Mehta, Rupeshkumar
Li, Jinyu
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
In video dubbing, aligning translated audio with the source audio is a significant challenge. Our focus is on achieving this efficiently, tailored for real-time, on-device video dubbing scenarios. We developed a phoneme-based end-to-end length-sensitive speech translation (LSST) model, which generates translations of varying lengths short, normal, and long using predefined tags. Additionally, we introduced length-aware beam search (LABS), an efficient approach to generate translations of different lengths in a single decoding pass. This approach maintained comparable BLEU scores compared to a baseline without length awareness while significantly enhancing synchronization quality between source and target audio, achieving a mean opinion score (MOS) gain of 0.34 for Spanish and 0.65 for Korean, respectively.
title Length Aware Speech Translation for Video Dubbing
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.00740