Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yaman, Dogucan, Akti, Seymanur, Eyiokur, Fevziye Irem, Waibel, Alexander
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914142770692096
author Yaman, Dogucan
Akti, Seymanur
Eyiokur, Fevziye Irem
Waibel, Alexander
author_facet Yaman, Dogucan
Akti, Seymanur
Eyiokur, Fevziye Irem
Waibel, Alexander
contents We propose a text-to-talking-face synthesis framework leveraging latent speech representations from HierSpeech++. A Text-to-Vec module generates Wav2Vec2 embeddings from text, which jointly condition speech and face generation. To handle distribution shifts between clean and TTS-predicted features, we adopt a two-stage training: pretraining on Wav2Vec2 embeddings and finetuning on TTS outputs. This enables tight audio-visual alignment, preserves speaker identity, and produces natural, expressive speech and synchronized facial motion without ground-truth audio at inference. Experiments show that conditioning on TTS-predicted latent features outperforms cascaded pipelines, improving both lip-sync and visual realism.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05432
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis
Yaman, Dogucan
Akti, Seymanur
Eyiokur, Fevziye Irem
Waibel, Alexander
Computer Vision and Pattern Recognition
We propose a text-to-talking-face synthesis framework leveraging latent speech representations from HierSpeech++. A Text-to-Vec module generates Wav2Vec2 embeddings from text, which jointly condition speech and face generation. To handle distribution shifts between clean and TTS-predicted features, we adopt a two-stage training: pretraining on Wav2Vec2 embeddings and finetuning on TTS outputs. This enables tight audio-visual alignment, preserves speaker identity, and produces natural, expressive speech and synchronized facial motion without ground-truth audio at inference. Experiments show that conditioning on TTS-predicted latent features outperforms cascaded pipelines, improving both lip-sync and visual realism.
title Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.05432