U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Xiang, Gao, Feng, Zhang, Yong, Pang, Youxin, Xiaoming, Xu, Kang, Zhuoliang, Wei, Xiaoming, Liu, Yebin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MAViD: A Multimodal Framework for Audio-Visual Dialogue Understanding and Generation
by: Pang, Youxin, et al.
Published: (2025)
by: Pang, Youxin, et al.
Published: (2025)
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
by: Pang, Youxin, et al.
Published: (2025)
by: Pang, Youxin, et al.
Published: (2025)
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
by: Kong, Zhe, et al.
Published: (2025)
by: Kong, Zhe, et al.
Published: (2025)
WildActor: Unconstrained Identity-Preserving Video Generation
by: Guo, Qin, et al.
Published: (2026)
by: Guo, Qin, et al.
Published: (2026)
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
by: Shao, Ruizhi, et al.
Published: (2024)
by: Shao, Ruizhi, et al.
Published: (2024)
Interpretable Perception and Reasoning for Audiovisual Geolocation
by: Su, Yiyang, et al.
Published: (2026)
by: Su, Yiyang, et al.
Published: (2026)
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
by: Yang, Shaoshu, et al.
Published: (2025)
by: Yang, Shaoshu, et al.
Published: (2025)
LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models
by: Yu, Haojie, et al.
Published: (2025)
by: Yu, Haojie, et al.
Published: (2025)
Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts
by: Deng, Xiang, et al.
Published: (2024)
by: Deng, Xiang, et al.
Published: (2024)
ARIG: Autoregressive Interactive Head Generation for Real-time Conversations
by: Guo, Ying, et al.
Published: (2025)
by: Guo, Ying, et al.
Published: (2025)
DAM-VSR: Disentanglement of Appearance and Motion for Video Super-Resolution
by: Kong, Zhe, et al.
Published: (2025)
by: Kong, Zhe, et al.
Published: (2025)
Active Intelligence in Video Avatars via Closed-loop World Modeling
by: He, Xuanhua, et al.
Published: (2025)
by: He, Xuanhua, et al.
Published: (2025)
ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping
by: Pang, Youxin, et al.
Published: (2024)
by: Pang, Youxin, et al.
Published: (2024)
Forge-and-Quench: Enhancing Image Generation for Higher Fidelity in Unified Multimodal Models
by: Zeng, Yanbing, et al.
Published: (2026)
by: Zeng, Yanbing, et al.
Published: (2026)
DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective
by: Chen, Yushuo, et al.
Published: (2025)
by: Chen, Yushuo, et al.
Published: (2025)
Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning
by: Liu, Xiaohong, et al.
Published: (2025)
by: Liu, Xiaohong, et al.
Published: (2025)
Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory
by: Wu, Ruiqi, et al.
Published: (2026)
by: Wu, Ruiqi, et al.
Published: (2026)
ECHO: Towards Emotionally Appropriate and Contextually Aware Interactive Head Generation
by: Kong, Xiangyu, et al.
Published: (2026)
by: Kong, Xiangyu, et al.
Published: (2026)
X-Streamer: Unified Human World Modeling with Audiovisual Interaction
by: Xie, You, et al.
Published: (2025)
by: Xie, You, et al.
Published: (2025)
PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework
by: Chen, SiXiang, et al.
Published: (2025)
by: Chen, SiXiang, et al.
Published: (2025)
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
by: Liu, Xiangyu, et al.
Published: (2026)
by: Liu, Xiangyu, et al.
Published: (2026)
MoGen: A Unified Collaborative Framework for Controllable Multi-Object Image Generation
by: Li, Yanfeng, et al.
Published: (2026)
by: Li, Yanfeng, et al.
Published: (2026)
BridgeNet: A Unified Multimodal Framework for Bridging 2D and 3D Industrial Anomaly Detection
by: Xiang, An, et al.
Published: (2025)
by: Xiang, An, et al.
Published: (2025)
PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback
by: Chen, Sixiang, et al.
Published: (2026)
by: Chen, Sixiang, et al.
Published: (2026)
X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction
by: Ren, Xiaoming, et al.
Published: (2026)
by: Ren, Xiaoming, et al.
Published: (2026)
HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene Perception
by: Yao, Wei, et al.
Published: (2025)
by: Yao, Wei, et al.
Published: (2025)
Generating Adversarial Events: A Motion-Aware Point Cloud Framework
by: Ren, Hongwei, et al.
Published: (2026)
by: Ren, Hongwei, et al.
Published: (2026)
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
by: Lee, Yebin, et al.
Published: (2024)
by: Lee, Yebin, et al.
Published: (2024)
LongCat-Video Technical Report
by: Meituan LongCat Team, et al.
Published: (2025)
by: Meituan LongCat Team, et al.
Published: (2025)
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
by: Tian, Rui, et al.
Published: (2025)
by: Tian, Rui, et al.
Published: (2025)
BEM: Balanced and Entropy-based Mix for Long-Tailed Semi-Supervised Learning
by: Zheng, Hongwei, et al.
Published: (2024)
by: Zheng, Hongwei, et al.
Published: (2024)
Elastic Interaction Energy-Informed Real-Time Traffic Scene Perception
by: Feng, Yaxin, et al.
Published: (2023)
by: Feng, Yaxin, et al.
Published: (2023)
LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation
by: Chern, Ethan, et al.
Published: (2025)
by: Chern, Ethan, et al.
Published: (2025)
Omni-Dish: Photorealistic and Faithful Image Generation and Editing for Arbitrary Chinese Dishes
by: Liu, Huijie, et al.
Published: (2025)
by: Liu, Huijie, et al.
Published: (2025)
SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios
by: Dang, Lingwei, et al.
Published: (2025)
by: Dang, Lingwei, et al.
Published: (2025)
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
by: Zhou, Donghao, et al.
Published: (2026)
by: Zhou, Donghao, et al.
Published: (2026)
PositionIC: Unified Position and Identity Consistency for Image Customization
by: Hu, Junjie, et al.
Published: (2025)
by: Hu, Junjie, et al.
Published: (2025)
EigenActor: Variant Body-Object Interaction Generation Evolved from Invariant Action Basis Reasoning
by: Gao, Xuehao, et al.
Published: (2025)
by: Gao, Xuehao, et al.
Published: (2025)
Similar Items
-
MAViD: A Multimodal Framework for Audio-Visual Dialogue Understanding and Generation
by: Pang, Youxin, et al.
Published: (2025) -
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
by: Pang, Youxin, et al.
Published: (2025) -
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
by: Kong, Zhe, et al.
Published: (2025) -
WildActor: Unconstrained Identity-Preserving Video Generation
by: Guo, Qin, et al.
Published: (2026) -
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
by: Shao, Ruizhi, et al.
Published: (2024)