Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Goncalves, Lucas, Mathur, Prashant, Niu, Xing, Houston, Brady, Lavania, Chandrashekhar, Vishnubhotla, Srikanth, Sun, Lijia, Ferritto, Anthony
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909437770334208
author Goncalves, Lucas
Mathur, Prashant
Niu, Xing
Houston, Brady
Lavania, Chandrashekhar
Vishnubhotla, Srikanth
Sun, Lijia
Ferritto, Anthony
author_facet Goncalves, Lucas
Mathur, Prashant
Niu, Xing
Houston, Brady
Lavania, Chandrashekhar
Vishnubhotla, Srikanth
Sun, Lijia
Ferritto, Anthony
contents Audio-Visual Speech-to-Speech Translation typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony-ensuring that the movements of the lips match the spoken content-essential for maintaining realism in dubbed videos. Despite its importance, the inclusion of lip-synchrony constraints in AVS2S models has been largely overlooked. This study addresses this gap by integrating a lip-synchrony loss into the training process of AVS2S models. Our proposed method significantly enhances lip-synchrony in direct audio-visual speech-to-speech translation, achieving an average LSE-D score of 10.67, representing a 9.2% reduction in LSE-D over a strong baseline across four language pairs. Additionally, it maintains the naturalness and high quality of the translated speech when overlaid onto the original video, without any degradation in translation quality.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16530
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
Goncalves, Lucas
Mathur, Prashant
Niu, Xing
Houston, Brady
Lavania, Chandrashekhar
Vishnubhotla, Srikanth
Sun, Lijia
Ferritto, Anthony
Sound
Computation and Language
Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
Audio-Visual Speech-to-Speech Translation typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony-ensuring that the movements of the lips match the spoken content-essential for maintaining realism in dubbed videos. Despite its importance, the inclusion of lip-synchrony constraints in AVS2S models has been largely overlooked. This study addresses this gap by integrating a lip-synchrony loss into the training process of AVS2S models. Our proposed method significantly enhances lip-synchrony in direct audio-visual speech-to-speech translation, achieving an average LSE-D score of 10.67, representing a 9.2% reduction in LSE-D over a strong baseline across four language pairs. Additionally, it maintains the naturalness and high quality of the translated speech when overlaid onto the original video, without any degradation in translation quality.
title Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
topic Sound
Computation and Language
Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2412.16530