MuteSwap: Visual-informed Silent Video Identity Conversion

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Yifan, Fang, Yu, Lin, Zhouhan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911087785410560
author Liu, Yifan
Fang, Yu
Lin, Zhouhan
author_facet Liu, Yifan
Fang, Yu
Lin, Zhouhan
contents Conventional voice conversion modifies voice characteristics from a source speaker to a target speaker, relying on audio input from both sides. However, this process becomes infeasible when clean audio is unavailable, such as in silent videos or noisy environments. In this work, we focus on the task of Silent Face-based Voice Conversion (SFVC), which does voice conversion entirely from visual inputs. i.e., given images of a target speaker and a silent video of a source speaker containing lip motion, SFVC generates speech aligning the identity of the target speaker while preserving the speech content in the source silent video. As this task requires generating intelligible speech and converting identity using only visual cues, it is particularly challenging. To address this, we introduce MuteSwap, a novel framework that employs contrastive learning to align cross-modality identities and minimize mutual information to separate shared visual features. Experimental results show that MuteSwap achieves impressive performance in both speech synthesis and identity conversion, especially under noisy conditions where methods dependent on audio input fail to produce intelligible results, demonstrating both the effectiveness of our training approach and the feasibility of SFVC.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00498
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MuteSwap: Visual-informed Silent Video Identity Conversion
Liu, Yifan
Fang, Yu
Lin, Zhouhan
Sound
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Audio and Speech Processing
Conventional voice conversion modifies voice characteristics from a source speaker to a target speaker, relying on audio input from both sides. However, this process becomes infeasible when clean audio is unavailable, such as in silent videos or noisy environments. In this work, we focus on the task of Silent Face-based Voice Conversion (SFVC), which does voice conversion entirely from visual inputs. i.e., given images of a target speaker and a silent video of a source speaker containing lip motion, SFVC generates speech aligning the identity of the target speaker while preserving the speech content in the source silent video. As this task requires generating intelligible speech and converting identity using only visual cues, it is particularly challenging. To address this, we introduce MuteSwap, a novel framework that employs contrastive learning to align cross-modality identities and minimize mutual information to separate shared visual features. Experimental results show that MuteSwap achieves impressive performance in both speech synthesis and identity conversion, especially under noisy conditions where methods dependent on audio input fail to produce intelligible results, demonstrating both the effectiveness of our training approach and the feasibility of SFVC.
title MuteSwap: Visual-informed Silent Video Identity Conversion
topic Sound
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2507.00498