EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Tian, Wenjie, Zhao, Zhixian, Hu, Jingbin, Chen, Huakang, Liu, Haohe, Mu, Binshen, Xie, Lei
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908871313850368
author Tian, Wenjie
Zhao, Zhixian
Hu, Jingbin
Chen, Huakang
Liu, Haohe
Mu, Binshen
Xie, Lei
author_facet Tian, Wenjie
Zhao, Zhixian
Hu, Jingbin
Chen, Huakang
Liu, Haohe
Mu, Binshen
Xie, Lei
contents The evolution of Omni-Modal Large Language Models~(Omni-LLMs) has revolutionized human--computer interaction, enabling unified audio-visual perception and speech response. However, existing Omni-LLMs struggle with complex real-world scenarios, often leading to superficial understanding and contextually mismatched emotional responses. This issue is further intensified by Omni-LLM's Thinker-Talker architectures, which are implicitly connected through hidden states, leading to the loss of emotional details. In this work, we present EmoOmni, a unified framework for accurate understanding and expression in multimodal emotional dialogue. At its core, we introduce the emotional Chain-of-Thought~(E-CoT), which enforces a reasoning from fine-grained multimodal perception to textual response. Moreover, we explicitly treat E-CoT as high-level emotional instructions that guide the talker, enabling accurate emotional expression. Complementing the model, we construct EmoOmniPipe to obtain the real-world annotated dialogue data and establish a benchmark, EmoOmniEval, to facilitate systematic assessment of multimodal emotional dialogue task. Experiments show that EmoOmni-7B achieves comparable performance with Qwen3Omni-30B-A3B-Thinking under the same talker.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21900
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs
Tian, Wenjie
Zhao, Zhixian
Hu, Jingbin
Chen, Huakang
Liu, Haohe
Mu, Binshen
Xie, Lei
Sound
Audio and Speech Processing
The evolution of Omni-Modal Large Language Models~(Omni-LLMs) has revolutionized human--computer interaction, enabling unified audio-visual perception and speech response. However, existing Omni-LLMs struggle with complex real-world scenarios, often leading to superficial understanding and contextually mismatched emotional responses. This issue is further intensified by Omni-LLM's Thinker-Talker architectures, which are implicitly connected through hidden states, leading to the loss of emotional details. In this work, we present EmoOmni, a unified framework for accurate understanding and expression in multimodal emotional dialogue. At its core, we introduce the emotional Chain-of-Thought~(E-CoT), which enforces a reasoning from fine-grained multimodal perception to textual response. Moreover, we explicitly treat E-CoT as high-level emotional instructions that guide the talker, enabling accurate emotional expression. Complementing the model, we construct EmoOmniPipe to obtain the real-world annotated dialogue data and establish a benchmark, EmoOmniEval, to facilitate systematic assessment of multimodal emotional dialogue task. Experiments show that EmoOmni-7B achieves comparable performance with Qwen3Omni-30B-A3B-Thinking under the same talker.
title EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2602.21900