Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912557466386432 |
|---|---|
| author | Tan, Weiting Lian, Jiachen Inaguma, Hirofumi Tomasello, Paden Koehn, Philipp Ma, Xutai |
| author_facet | Tan, Weiting Lian, Jiachen Inaguma, Hirofumi Tomasello, Paden Koehn, Philipp Ma, Xutai |
| contents | We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion strategies during pre-training to identify the most effective integration approach. Subsequent fine-tuning on emotion recognition and expressive dialogue tasks yields substantial gains over speech-only baselines (e.g., +5 F1 in emotion recognition). AVLM highlights the value of expressive visual information in guiding speech generation and offers a foundation for end-to-end multimodal conversational systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_16188 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation Tan, Weiting Lian, Jiachen Inaguma, Hirofumi Tomasello, Paden Koehn, Philipp Ma, Xutai Computation and Language Computer Vision and Pattern Recognition Multimedia Sound Audio and Speech Processing We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion strategies during pre-training to identify the most effective integration approach. Subsequent fine-tuning on emotion recognition and expressive dialogue tasks yields substantial gains over speech-only baselines (e.g., +5 F1 in emotion recognition). AVLM highlights the value of expressive visual information in guiding speech generation and offers a foundation for end-to-end multimodal conversational systems. |
| title | Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation |
| topic | Computation and Language Computer Vision and Pattern Recognition Multimedia Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2508.16188 |