Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Taesoo, Jo, Yongsik, Song, Hyunmin, Kim, Taehwan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Transformer Models for Turn-taking Prediction: Effects on Conversational Dynamics of Human-Agent Interaction during Cooperative Gameplay
by: Bae, Young-Ho, et al.
Published: (2025)
by: Bae, Young-Ho, et al.
Published: (2025)
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
by: Sung, Yoo Yeon, et al.
Published: (2025)
by: Sung, Yoo Yeon, et al.
Published: (2025)
OpenOmni: A Collaborative Open Source Tool for Building Future-Ready Multimodal Conversational Agents
by: Sun, Qiang, et al.
Published: (2024)
by: Sun, Qiang, et al.
Published: (2024)
Humane Speech Synthesis through Zero-Shot Emotion and Disfluency Generation
by: Chaudhury, Rohan, et al.
Published: (2024)
by: Chaudhury, Rohan, et al.
Published: (2024)
Dyadic: A Scalable Platform for Human-Human and Human-AI Conversation Research
by: Markowitz, David M.
Published: (2026)
by: Markowitz, David M.
Published: (2026)
ReSpAct: Harmonizing Reasoning, Speaking, and Acting Towards Building Large Language Model-Based Conversational AI Agents
by: Dongre, Vardhan, et al.
Published: (2024)
by: Dongre, Vardhan, et al.
Published: (2024)
Human Latency Conversational Turns for Spoken Avatar Systems
by: Jacoby, Derek, et al.
Published: (2024)
by: Jacoby, Derek, et al.
Published: (2024)
Empathic Prompting: Non-Verbal Context Integration for Multimodal LLM Conversations
by: Stacchio, Lorenzo, et al.
Published: (2025)
by: Stacchio, Lorenzo, et al.
Published: (2025)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes
by: Chen, Maximillian, et al.
Published: (2026)
by: Chen, Maximillian, et al.
Published: (2026)
Beyond Preset Identities: How Agents Form Stances and Boundaries in Generative Societies
by: Zhang, Hanzhong, et al.
Published: (2026)
by: Zhang, Hanzhong, et al.
Published: (2026)
One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations
by: Lee, Yoonjoo, et al.
Published: (2024)
by: Lee, Yoonjoo, et al.
Published: (2024)
Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
by: Yao, Bingsheng, et al.
Published: (2025)
by: Yao, Bingsheng, et al.
Published: (2025)
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
by: Yoon, Sion, et al.
Published: (2024)
by: Yoon, Sion, et al.
Published: (2024)
You Only Look at Screens: Multimodal Chain-of-Action Agents
by: Zhang, Zhuosheng, et al.
Published: (2023)
by: Zhang, Zhuosheng, et al.
Published: (2023)
Human-AI Collaborative Taxonomy Construction: A Case Study in Profession-Specific Writing Assistants
by: Lee, Minhwa, et al.
Published: (2024)
by: Lee, Minhwa, et al.
Published: (2024)
Position: Towards Bidirectional Human-AI Alignment
by: Shen, Hua, et al.
Published: (2024)
by: Shen, Hua, et al.
Published: (2024)
Explore, Select, Derive, and Recall: Augmenting LLM with Human-like Memory for Mobile Task Automation
by: Lee, Sunjae, et al.
Published: (2023)
by: Lee, Sunjae, et al.
Published: (2023)
Human and LLM Biases in Hate Speech Annotations: A Socio-Demographic Analysis of Annotators and Targets
by: Giorgi, Tommaso, et al.
Published: (2024)
by: Giorgi, Tommaso, et al.
Published: (2024)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
by: Kim, Jane Paik
Published: (2026)
by: Kim, Jane Paik
Published: (2026)
Role-Play Zero-Shot Prompting with Large Language Models for Open-Domain Human-Machine Conversation
by: Njifenjou, Ahmed, et al.
Published: (2024)
by: Njifenjou, Ahmed, et al.
Published: (2024)
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
by: Oh, Juhyun, et al.
Published: (2024)
by: Oh, Juhyun, et al.
Published: (2024)
MASCOT: Towards Multi-Agent Socio-Collaborative Companion Systems
by: Wang, Yiyang, et al.
Published: (2026)
by: Wang, Yiyang, et al.
Published: (2026)
Autonomous Workflow for Multimodal Fine-Grained Training Assistants Towards Mixed Reality
by: Pei, Jiahuan, et al.
Published: (2024)
by: Pei, Jiahuan, et al.
Published: (2024)
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
by: Kim, Tae Soo, et al.
Published: (2023)
by: Kim, Tae Soo, et al.
Published: (2023)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
by: Wang, Zora Zhiruo, et al.
Published: (2025)
by: Wang, Zora Zhiruo, et al.
Published: (2025)
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
by: Shao, Yijia, et al.
Published: (2024)
by: Shao, Yijia, et al.
Published: (2024)
Can Large Language Model Agents Simulate Human Trust Behavior?
by: Xie, Chengxing, et al.
Published: (2024)
by: Xie, Chengxing, et al.
Published: (2024)
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
by: Wu, Zongru, et al.
Published: (2025)
by: Wu, Zongru, et al.
Published: (2025)
VizTrust: A Visual Analytics Tool for Capturing User Trust Dynamics in Human-AI Communication
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback
by: Lee, Dong Won, et al.
Published: (2024)
by: Lee, Dong Won, et al.
Published: (2024)
Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks
by: Rahman, Hasibur, et al.
Published: (2025)
by: Rahman, Hasibur, et al.
Published: (2025)
From Human-to-Human to Human-to-Bot Conversations in Software Engineering
by: Khojah, Ranim, et al.
Published: (2024)
by: Khojah, Ranim, et al.
Published: (2024)
Keeping Users Engaged During Repeated Administration of the Same Questionnaire: Using Large Language Models to Reliably Diversify Questions
by: Yun, Hye Sun, et al.
Published: (2023)
by: Yun, Hye Sun, et al.
Published: (2023)
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
by: Huq, Faria, et al.
Published: (2025)
by: Huq, Faria, et al.
Published: (2025)
From Control to Foresight: Simulation as a New Paradigm for Human-Agent Collaboration
by: He, Gaole, et al.
Published: (2026)
by: He, Gaole, et al.
Published: (2026)
Causal Autoencoder-like Generation of Feedback Fuzzy Cognitive Maps with an LLM Agent
by: Panda, Akash Kumar, et al.
Published: (2025)
by: Panda, Akash Kumar, et al.
Published: (2025)
Similar Items
-
Multimodal Transformer Models for Turn-taking Prediction: Effects on Conversational Dynamics of Human-Agent Interaction during Cooperative Gameplay
by: Bae, Young-Ho, et al.
Published: (2025) -
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
by: Sung, Yoo Yeon, et al.
Published: (2025) -
OpenOmni: A Collaborative Open Source Tool for Building Future-Ready Multimodal Conversational Agents
by: Sun, Qiang, et al.
Published: (2024) -
Humane Speech Synthesis through Zero-Shot Emotion and Disfluency Generation
by: Chaudhury, Rohan, et al.
Published: (2024) -
Dyadic: A Scalable Platform for Human-Human and Human-AI Conversation Research
by: Markowitz, David M.
Published: (2026)