U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917299263373312 |
|---|---|
| author | Deng, Xiang Gao, Feng Zhang, Yong Pang, Youxin Xiaoming, Xu Kang, Zhuoliang Wei, Xiaoming Liu, Yebin |
| author_facet | Deng, Xiang Gao, Feng Zhang, Yong Pang, Youxin Xiaoming, Xu Kang, Zhuoliang Wei, Xiaoming Liu, Yebin |
| contents | Full-stack multimodal interaction in real-time is a central goal in building intelligent embodied agents capable of natural, dynamic communication. However, existing systems are either limited to unimodal generation or suffer from degraded reasoning and poor cross-modal alignment, preventing coherent and perceptually grounded interactions. In this work, we introduce U-Mind, the first unified system for high-intelligence multimodal dialogue that supports real-time generation and jointly models language, speech, motion, and video synthesis within a single interactive loop. At its core, U-Mind implements a Unified Alignment and Reasoning Framework that addresses two key challenges: enhancing cross-modal synchronization via a segment-wise alignment strategy, and preserving reasoning abilities through Rehearsal-Driven Learning. During inference, U-Mind adopts a text-first decoding pipeline that performs internal chain-of-thought planning followed by temporally synchronized generation across modalities. To close the loop, we implement a real-time video rendering framework conditioned on pose and speech, enabling expressive and synchronized visual feedback. Extensive experiments demonstrate that U-Mind achieves state-of-the-art performance on a range of multimodal interaction tasks, including question answering, instruction following, and motion generation, paving the way toward intelligent, immersive conversational agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_23739 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation Deng, Xiang Gao, Feng Zhang, Yong Pang, Youxin Xiaoming, Xu Kang, Zhuoliang Wei, Xiaoming Liu, Yebin Computer Vision and Pattern Recognition Full-stack multimodal interaction in real-time is a central goal in building intelligent embodied agents capable of natural, dynamic communication. However, existing systems are either limited to unimodal generation or suffer from degraded reasoning and poor cross-modal alignment, preventing coherent and perceptually grounded interactions. In this work, we introduce U-Mind, the first unified system for high-intelligence multimodal dialogue that supports real-time generation and jointly models language, speech, motion, and video synthesis within a single interactive loop. At its core, U-Mind implements a Unified Alignment and Reasoning Framework that addresses two key challenges: enhancing cross-modal synchronization via a segment-wise alignment strategy, and preserving reasoning abilities through Rehearsal-Driven Learning. During inference, U-Mind adopts a text-first decoding pipeline that performs internal chain-of-thought planning followed by temporally synchronized generation across modalities. To close the loop, we implement a real-time video rendering framework conditioned on pose and speech, enabling expressive and synchronized visual feedback. Extensive experiments demonstrate that U-Mind achieves state-of-the-art performance on a range of multimodal interaction tasks, including question answering, instruction following, and motion generation, paving the way toward intelligent, immersive conversational agents. |
| title | U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2602.23739 |