UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Pang, Youxin, Zhang, Yong, Shao, Ruizhi, Deng, Xiang, Gao, Feng, Xiaoming, Xu, Wei, Xiaoming, Liu, Yebin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation
by: Deng, Xiang, et al.
Published: (2026)
by: Deng, Xiang, et al.
Published: (2026)
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
by: Shao, Ruizhi, et al.
Published: (2024)
by: Shao, Ruizhi, et al.
Published: (2024)
MAViD: A Multimodal Framework for Audio-Visual Dialogue Understanding and Generation
by: Pang, Youxin, et al.
Published: (2025)
by: Pang, Youxin, et al.
Published: (2025)
UniMo: Unified Motion Generation and Understanding with Chain of Thought
by: Wang, Guocun, et al.
Published: (2026)
by: Wang, Guocun, et al.
Published: (2026)
DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective
by: Chen, Yushuo, et al.
Published: (2025)
by: Chen, Yushuo, et al.
Published: (2025)
SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios
by: Dang, Lingwei, et al.
Published: (2025)
by: Dang, Lingwei, et al.
Published: (2025)
ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping
by: Pang, Youxin, et al.
Published: (2024)
by: Pang, Youxin, et al.
Published: (2024)
Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts
by: Deng, Xiang, et al.
Published: (2024)
by: Deng, Xiang, et al.
Published: (2024)
Interspatial Attention for Efficient 4D Human Video Generation
by: Shao, Ruizhi, et al.
Published: (2025)
by: Shao, Ruizhi, et al.
Published: (2025)
The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion
by: Chen, Changan, et al.
Published: (2024)
by: Chen, Changan, et al.
Published: (2024)
Layered 3D Human Generation via Semantic-Aware Diffusion Model
by: Wang, Yi, et al.
Published: (2023)
by: Wang, Yi, et al.
Published: (2023)
Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation
by: Cao, Chenjie, et al.
Published: (2025)
by: Cao, Chenjie, et al.
Published: (2025)
H-MoRe: Learning Human-centric Motion Representation for Action Analysis
by: Huang, Zhanbo, et al.
Published: (2025)
by: Huang, Zhanbo, et al.
Published: (2025)
Uni4D: A Unified Self-Supervised Learning Framework for Point Cloud Videos
by: Zuo, Zhi, et al.
Published: (2025)
by: Zuo, Zhi, et al.
Published: (2025)
GPS-Gaussian: Generalizable Pixel-wise 3D Gaussian Splatting for Real-time Human Novel View Synthesis
by: Zheng, Shunyuan, et al.
Published: (2023)
by: Zheng, Shunyuan, et al.
Published: (2023)
Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos
by: Gong, Kehong, et al.
Published: (2025)
by: Gong, Kehong, et al.
Published: (2025)
GPS-Gaussian+: Generalizable Pixel-wise 3D Gaussian Splatting for Real-Time Human-Scene Rendering from Sparse Views
by: Zhou, Boyao, et al.
Published: (2024)
by: Zhou, Boyao, et al.
Published: (2024)
UniVision: A Unified Framework for Vision-Centric 3D Perception
by: Hong, Yu, et al.
Published: (2024)
by: Hong, Yu, et al.
Published: (2024)
GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
by: Chai, Ying, et al.
Published: (2025)
by: Chai, Ying, et al.
Published: (2025)
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths
by: Mao, Weijia, et al.
Published: (2025)
by: Mao, Weijia, et al.
Published: (2025)
UniCon3R: Unified Contact-aware 4D Human-Scene Reconstruction from Monocular Video
by: Sur, Tanuj, et al.
Published: (2026)
by: Sur, Tanuj, et al.
Published: (2026)
CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos
by: Zhao, Chengfeng, et al.
Published: (2026)
by: Zhao, Chengfeng, et al.
Published: (2026)
UniCorrn: Unified Correspondence Transformer Across 2D and 3D
by: Goswami, Prajnan, et al.
Published: (2026)
by: Goswami, Prajnan, et al.
Published: (2026)
UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing
by: Bai, Jianhong, et al.
Published: (2024)
by: Bai, Jianhong, et al.
Published: (2024)
UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation
by: Wang, Xiang, et al.
Published: (2024)
by: Wang, Xiang, et al.
Published: (2024)
DreamCraft3D++: Efficient Hierarchical 3D Generation with Multi-Plane Reconstruction Model
by: Sun, Jingxiang, et al.
Published: (2024)
by: Sun, Jingxiang, et al.
Published: (2024)
HumanCoser: Layered 3D Human Generation via Semantic-Aware Diffusion Model
by: Wang, Yi, et al.
Published: (2024)
by: Wang, Yi, et al.
Published: (2024)
PnP-U3D: Plug-and-Play 3D Framework Bridging Autoregression and Diffusion for Unified Understanding and Generation
by: Chen, Yongwei, et al.
Published: (2026)
by: Chen, Yongwei, et al.
Published: (2026)
4DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular Video
by: Lyu, Jin, et al.
Published: (2026)
by: Lyu, Jin, et al.
Published: (2026)
MOSS: Motion-based 3D Clothed Human Synthesis from Monocular Video
by: Wang, Hongsheng, et al.
Published: (2024)
by: Wang, Hongsheng, et al.
Published: (2024)
UniMuMo: Unified Text, Music and Motion Generation
by: Yang, Han, et al.
Published: (2024)
by: Yang, Han, et al.
Published: (2024)
DAM-VSR: Disentanglement of Appearance and Motion for Video Super-Resolution
by: Kong, Zhe, et al.
Published: (2025)
by: Kong, Zhe, et al.
Published: (2025)
UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation
by: Zhang, Ruiheng, et al.
Published: (2026)
by: Zhang, Ruiheng, et al.
Published: (2026)
MoSa: Motion Generation with Scalable Autoregressive Modeling
by: Liu, Mengyuan, et al.
Published: (2025)
by: Liu, Mengyuan, et al.
Published: (2025)
A Unified Framework for 3D Scene Understanding
by: Xu, Wei, et al.
Published: (2024)
by: Xu, Wei, et al.
Published: (2024)
UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Functionality Segmentation
by: Lin, Jiaying, et al.
Published: (2026)
by: Lin, Jiaying, et al.
Published: (2026)
UniT: Unified Geometry Learning with Group Autoregressive Transformer
by: Wang, Haotian, et al.
Published: (2026)
by: Wang, Haotian, et al.
Published: (2026)
TRIM: Scalable 3D Gaussian Diffusion Inference with Temporal and Spatial Trimming
by: Yin, Zeyuan, et al.
Published: (2025)
by: Yin, Zeyuan, et al.
Published: (2025)
X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents
by: Song, Guoxian, et al.
Published: (2025)
by: Song, Guoxian, et al.
Published: (2025)
Similar Items
-
U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation
by: Deng, Xiang, et al.
Published: (2026) -
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
by: Shao, Ruizhi, et al.
Published: (2024) -
MAViD: A Multimodal Framework for Audio-Visual Dialogue Understanding and Generation
by: Pang, Youxin, et al.
Published: (2025) -
UniMo: Unified Motion Generation and Understanding with Chain of Thought
by: Wang, Guocun, et al.
Published: (2026) -
DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual Perspective
by: Chen, Yushuo, et al.
Published: (2025)