Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tan, Weiting, Lian, Jiachen, Inaguma, Hirofumi, Tomasello, Paden, Koehn, Philipp, Ma, Xutai
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912557466386432
author Tan, Weiting
Lian, Jiachen
Inaguma, Hirofumi
Tomasello, Paden
Koehn, Philipp
Ma, Xutai
author_facet Tan, Weiting
Lian, Jiachen
Inaguma, Hirofumi
Tomasello, Paden
Koehn, Philipp
Ma, Xutai
contents We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion strategies during pre-training to identify the most effective integration approach. Subsequent fine-tuning on emotion recognition and expressive dialogue tasks yields substantial gains over speech-only baselines (e.g., +5 F1 in emotion recognition). AVLM highlights the value of expressive visual information in guiding speech generation and offers a foundation for end-to-end multimodal conversational systems.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16188
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
Tan, Weiting
Lian, Jiachen
Inaguma, Hirofumi
Tomasello, Paden
Koehn, Philipp
Ma, Xutai
Computation and Language
Computer Vision and Pattern Recognition
Multimedia
Sound
Audio and Speech Processing
We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion strategies during pre-training to identify the most effective integration approach. Subsequent fine-tuning on emotion recognition and expressive dialogue tasks yields substantial gains over speech-only baselines (e.g., +5 F1 in emotion recognition). AVLM highlights the value of expressive visual information in guiding speech generation and offers a foundation for end-to-end multimodal conversational systems.
title Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
topic Computation and Language
Computer Vision and Pattern Recognition
Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2508.16188