Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866914142770692096 |
|---|---|
| author | Yaman, Dogucan Akti, Seymanur Eyiokur, Fevziye Irem Waibel, Alexander |
| author_facet | Yaman, Dogucan Akti, Seymanur Eyiokur, Fevziye Irem Waibel, Alexander |
| contents | We propose a text-to-talking-face synthesis framework leveraging latent speech representations from HierSpeech++. A Text-to-Vec module generates Wav2Vec2 embeddings from text, which jointly condition speech and face generation. To handle distribution shifts between clean and TTS-predicted features, we adopt a two-stage training: pretraining on Wav2Vec2 embeddings and finetuning on TTS outputs. This enables tight audio-visual alignment, preserves speaker identity, and produces natural, expressive speech and synchronized facial motion without ground-truth audio at inference. Experiments show that conditioning on TTS-predicted latent features outperforms cascaded pipelines, improving both lip-sync and visual realism. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_05432 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis Yaman, Dogucan Akti, Seymanur Eyiokur, Fevziye Irem Waibel, Alexander Computer Vision and Pattern Recognition We propose a text-to-talking-face synthesis framework leveraging latent speech representations from HierSpeech++. A Text-to-Vec module generates Wav2Vec2 embeddings from text, which jointly condition speech and face generation. To handle distribution shifts between clean and TTS-predicted features, we adopt a two-stage training: pretraining on Wav2Vec2 embeddings and finetuning on TTS outputs. This enables tight audio-visual alignment, preserves speaker identity, and produces natural, expressive speech and synchronized facial motion without ground-truth audio at inference. Experiments show that conditioning on TTS-predicted latent features outperforms cascaded pipelines, improving both lip-sync and visual realism. |
| title | Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.05432 |