ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913052094365696 |
|---|---|
| author | Sun, Zhengwentai Zheng, Keru Li, Chenghong Liao, Hongjie Yang, Xihe Li, Heyuan Zhi, Yihao Ning, Shuliang Cui, Shuguang Han, Xiaoguang |
| author_facet | Sun, Zhengwentai Zheng, Keru Li, Chenghong Liao, Hongjie Yang, Xihe Li, Heyuan Zhi, Yihao Ning, Shuliang Cui, Shuguang Han, Xiaoguang |
| contents | Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi-view data. Existing methods often address these factors separately, resulting in limited controllability or reduced visual quality. We revisit this problem from an image-first perspective, where high-quality human appearance is learned via image generation and used as a prior for video synthesis, decoupling appearance modeling from temporal consistency. We propose a pose- and viewpoint-controllable pipeline that combines a pretrained image backbone with SMPL-X-based motion guidance, together with a training-free temporal refinement stage based on a pretrained video diffusion model. Our method produces high-quality, temporally consistent videos under diverse poses and viewpoints. We also release a canonical human dataset and an auxiliary model for compositional human image synthesis. Code and data are publicly available at https://github.com/Taited/ReImagine. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_19720 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis Sun, Zhengwentai Zheng, Keru Li, Chenghong Liao, Hongjie Yang, Xihe Li, Heyuan Zhi, Yihao Ning, Shuliang Cui, Shuguang Han, Xiaoguang Computer Vision and Pattern Recognition Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi-view data. Existing methods often address these factors separately, resulting in limited controllability or reduced visual quality. We revisit this problem from an image-first perspective, where high-quality human appearance is learned via image generation and used as a prior for video synthesis, decoupling appearance modeling from temporal consistency. We propose a pose- and viewpoint-controllable pipeline that combines a pretrained image backbone with SMPL-X-based motion guidance, together with a training-free temporal refinement stage based on a pretrained video diffusion model. Our method produces high-quality, temporally consistent videos under diverse poses and viewpoints. We also release a canonical human dataset and an auxiliary model for compositional human image synthesis. Code and data are publicly available at https://github.com/Taited/ReImagine. |
| title | ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2604.19720 |