MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zheng, Junjie, Chen, Zihao, Ding, Chaofan, Liang, Yunming, Fan, Yihan, Yang, Huan, Xie, Lei, Di, Xinhan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912387019309056
author Zheng, Junjie
Chen, Zihao
Ding, Chaofan
Liang, Yunming
Fan, Yihan
Yang, Huan
Xie, Lei
Di, Xinhan
author_facet Zheng, Junjie
Chen, Zihao
Ding, Chaofan
Liang, Yunming
Fan, Yihan
Yang, Huan
Xie, Lei
Di, Xinhan
contents Current movie dubbing technology can produce the desired speech using a reference voice and input video, maintaining perfect synchronization with the visuals while effectively conveying the intended emotions. However, crucial aspects of movie dubbing, including adaptation to various dubbing styles, effective handling of dialogue, narration, and monologues, as well as consideration of subtle details such as speaker age and gender, remain insufficiently explored. To tackle these challenges, we introduce a multi-modal generative framework. First, it utilizes a multi-modal large vision-language model (VLM) to analyze visual inputs, enabling the recognition of dubbing types and fine-grained attributes. Second, it produces high-quality dubbing using large speech generation models, guided by multi-modal inputs. Additionally, a movie dubbing dataset with annotations for dubbing types and subtle details is constructed to enhance movie understanding and improve dubbing quality for the proposed multi-modal framework. Experimental results across multiple benchmark datasets show superior performance compared to state-of-the-art (SOTA) methods. In details, the LSE-D, SPK-SIM, EMO-SIM, and MCD exhibit improvements of up to 1.09%, 8.80%, 19.08%, and 18.74%, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16279
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
Zheng, Junjie
Chen, Zihao
Ding, Chaofan
Liang, Yunming
Fan, Yihan
Yang, Huan
Xie, Lei
Di, Xinhan
Multimedia
Computer Vision and Pattern Recognition
Current movie dubbing technology can produce the desired speech using a reference voice and input video, maintaining perfect synchronization with the visuals while effectively conveying the intended emotions. However, crucial aspects of movie dubbing, including adaptation to various dubbing styles, effective handling of dialogue, narration, and monologues, as well as consideration of subtle details such as speaker age and gender, remain insufficiently explored. To tackle these challenges, we introduce a multi-modal generative framework. First, it utilizes a multi-modal large vision-language model (VLM) to analyze visual inputs, enabling the recognition of dubbing types and fine-grained attributes. Second, it produces high-quality dubbing using large speech generation models, guided by multi-modal inputs. Additionally, a movie dubbing dataset with annotations for dubbing types and subtle details is constructed to enhance movie understanding and improve dubbing quality for the proposed multi-modal framework. Experimental results across multiple benchmark datasets show superior performance compared to state-of-the-art (SOTA) methods. In details, the LSE-D, SPK-SIM, EMO-SIM, and MCD exhibit improvements of up to 1.09%, 8.80%, 19.08%, and 18.74%, respectively.
title MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
topic Multimedia
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.16279