Beyond Words: Multimodal LLM Knows When to Speak

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liao, Zikai, Ouyang, Yi, Lee, Yi-Lun, Yu, Chen-Ping, Tsai, Yi-Hsuan, Yin, Zhaozheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910237589504000
author Liao, Zikai
Ouyang, Yi
Lee, Yi-Lun
Yu, Chen-Ping
Tsai, Yi-Hsuan
Yin, Zhaozheng
author_facet Liao, Zikai
Ouyang, Yi
Lee, Yi-Lun
Yu, Chen-Ping
Tsai, Yi-Hsuan
Yin, Zhaozheng
contents Chatbots via large language models (LLMs) generate fluent responses but often struggle with when to speak, especially for brief, timely listener reactions during ongoing dialogue. We present a multimodal strategy for LLMs, which leverages synchronized video, audio, and text cues to improve conversational timing awareness. The strategy reformulates response timing as a dense response-type prediction task, enabling an agent to decide whether to remain silent, produce a short reaction, or start a full response under streaming constraints. Therefore, we introduce a curated multimodal dataset from real-world dyadic conversational videos with temporally aligned modalities and fine-grained reaction type annotations. Moreover, we design a multimodal strategy, MM-When2Speak, with a multimodal integration module on top of an LLM backbone. Experiments across various modality settings and strong LLM baselines show that MM-When2Speak achieves up to a 3x improvement in response type prediction performance, highlighting the importance of multimodal perception for natural and engaging conversational interaction.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14654
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Words: Multimodal LLM Knows When to Speak
Liao, Zikai
Ouyang, Yi
Lee, Yi-Lun
Yu, Chen-Ping
Tsai, Yi-Hsuan
Yin, Zhaozheng
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Chatbots via large language models (LLMs) generate fluent responses but often struggle with when to speak, especially for brief, timely listener reactions during ongoing dialogue. We present a multimodal strategy for LLMs, which leverages synchronized video, audio, and text cues to improve conversational timing awareness. The strategy reformulates response timing as a dense response-type prediction task, enabling an agent to decide whether to remain silent, produce a short reaction, or start a full response under streaming constraints. Therefore, we introduce a curated multimodal dataset from real-world dyadic conversational videos with temporally aligned modalities and fine-grained reaction type annotations. Moreover, we design a multimodal strategy, MM-When2Speak, with a multimodal integration module on top of an LLM backbone. Experiments across various modality settings and strong LLM baselines show that MM-When2Speak achieves up to a 3x improvement in response type prediction performance, highlighting the importance of multimodal perception for natural and engaging conversational interaction.
title Beyond Words: Multimodal LLM Knows When to Speak
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.14654