Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kim, Taesoo, Jo, Yongsik, Song, Hyunmin, Kim, Taehwan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912591484289024
author Kim, Taesoo
Jo, Yongsik
Song, Hyunmin
Kim, Taehwan
author_facet Kim, Taesoo
Jo, Yongsik
Song, Hyunmin
Kim, Taehwan
contents Human conversation involves language, speech, and visual cues, with each medium providing complementary information. For instance, speech conveys a vibe or tone not fully captured by text alone. While multimodal LLMs focus on generating text responses from diverse inputs, less attention has been paid to generating natural and engaging speech. We propose a human-like agent that generates speech responses based on conversation mood and responsive style information. To achieve this, we build a novel MultiSensory Conversation dataset focused on speech to enable agents to generate natural speech. We then propose a multimodal LLM-based model for generating text responses and voice descriptions, which are used to generate speech covering paralinguistic information. Experimental results demonstrate the effectiveness of utilizing both visual and audio modalities in conversation to generate engaging speech. The source code is available in https://github.com/kimtaesu24/MSenC
format Preprint
id arxiv_https___arxiv_org_abs_2509_14627
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
Kim, Taesoo
Jo, Yongsik
Song, Hyunmin
Kim, Taehwan
Human-Computer Interaction
Artificial Intelligence
Computation and Language
Human conversation involves language, speech, and visual cues, with each medium providing complementary information. For instance, speech conveys a vibe or tone not fully captured by text alone. While multimodal LLMs focus on generating text responses from diverse inputs, less attention has been paid to generating natural and engaging speech. We propose a human-like agent that generates speech responses based on conversation mood and responsive style information. To achieve this, we build a novel MultiSensory Conversation dataset focused on speech to enable agents to generate natural speech. We then propose a multimodal LLM-based model for generating text responses and voice descriptions, which are used to generate speech covering paralinguistic information. Experimental results demonstrate the effectiveness of utilizing both visual and audio modalities in conversation to generate engaging speech. The source code is available in https://github.com/kimtaesu24/MSenC
title Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
topic Human-Computer Interaction
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.14627