Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhao, Chengfeng, Shu, Jiazhi, Zhao, Yubo, Huang, Tianyu, Lu, Jiahao, Gu, Zekai, Ren, Chengwei, Dou, Zhiyang, Shuai, Qing, Liu, Yuan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2601.10632
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914463740854272
author Zhao, Chengfeng
Shu, Jiazhi
Zhao, Yubo
Huang, Tianyu
Lu, Jiahao
Gu, Zekai
Ren, Chengwei
Dou, Zhiyang
Shuai, Qing
Liu, Yuan
author_facet Zhao, Chengfeng
Shu, Jiazhi
Zhao, Yubo
Huang, Tianyu
Lu, Jiahao
Gu, Zekai
Ren, Chengwei
Dou, Zhiyang
Shuai, Qing
Liu, Yuan
contents In this paper, we find that the generation of 3D human motions and 2D human videos is intrinsically coupled. 3D motions provide the structural prior for plausibility and consistency in videos, while pre-trained video models offer strong generalization capabilities for motions. Based on this, we present CoMoVi, a co-generative framework that generates 3D human motions and videos synchronously within a single diffusion denoising loop. However, since the 3D human motions and the 2D human-centric videos have a modality gap between each other, we propose to project the 3D human motion into an effective 2D human motion representation that effectively aligns with the 2D videos. Then, we design a dual-branch diffusion model to couple human motion and the video generation process with mutual feature interaction and 3D-2D cross attentions. To train and evaluate our model, we curate CoMoVi-Dataset, a large-scale real-world human video dataset with text and motion annotations, covering diverse and challenging human motions. Extensive experiments demonstrate that our method generates high-quality 3D human motion with a better generalization ability and that our method can generate high-quality human-centric videos without external motion references.
format Preprint
id arxiv_https___arxiv_org_abs_2601_10632
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos
Zhao, Chengfeng
Shu, Jiazhi
Zhao, Yubo
Huang, Tianyu
Lu, Jiahao
Gu, Zekai
Ren, Chengwei
Dou, Zhiyang
Shuai, Qing
Liu, Yuan
Computer Vision and Pattern Recognition
In this paper, we find that the generation of 3D human motions and 2D human videos is intrinsically coupled. 3D motions provide the structural prior for plausibility and consistency in videos, while pre-trained video models offer strong generalization capabilities for motions. Based on this, we present CoMoVi, a co-generative framework that generates 3D human motions and videos synchronously within a single diffusion denoising loop. However, since the 3D human motions and the 2D human-centric videos have a modality gap between each other, we propose to project the 3D human motion into an effective 2D human motion representation that effectively aligns with the 2D videos. Then, we design a dual-branch diffusion model to couple human motion and the video generation process with mutual feature interaction and 3D-2D cross attentions. To train and evaluate our model, we curate CoMoVi-Dataset, a large-scale real-world human video dataset with text and motion annotations, covering diverse and challenging human motions. Extensive experiments demonstrate that our method generates high-quality 3D human motion with a better generalization ability and that our method can generate high-quality human-centric videos without external motion references.
title CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.10632