ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Zhengwentai, Zheng, Keru, Li, Chenghong, Liao, Hongjie, Yang, Xihe, Li, Heyuan, Zhi, Yihao, Ning, Shuliang, Cui, Shuguang, Han, Xiaoguang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913052094365696
author Sun, Zhengwentai
Zheng, Keru
Li, Chenghong
Liao, Hongjie
Yang, Xihe
Li, Heyuan
Zhi, Yihao
Ning, Shuliang
Cui, Shuguang
Han, Xiaoguang
author_facet Sun, Zhengwentai
Zheng, Keru
Li, Chenghong
Liao, Hongjie
Yang, Xihe
Li, Heyuan
Zhi, Yihao
Ning, Shuliang
Cui, Shuguang
Han, Xiaoguang
contents Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi-view data. Existing methods often address these factors separately, resulting in limited controllability or reduced visual quality. We revisit this problem from an image-first perspective, where high-quality human appearance is learned via image generation and used as a prior for video synthesis, decoupling appearance modeling from temporal consistency. We propose a pose- and viewpoint-controllable pipeline that combines a pretrained image backbone with SMPL-X-based motion guidance, together with a training-free temporal refinement stage based on a pretrained video diffusion model. Our method produces high-quality, temporally consistent videos under diverse poses and viewpoints. We also release a canonical human dataset and an auxiliary model for compositional human image synthesis. Code and data are publicly available at https://github.com/Taited/ReImagine.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19720
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis
Sun, Zhengwentai
Zheng, Keru
Li, Chenghong
Liao, Hongjie
Yang, Xihe
Li, Heyuan
Zhi, Yihao
Ning, Shuliang
Cui, Shuguang
Han, Xiaoguang
Computer Vision and Pattern Recognition
Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi-view data. Existing methods often address these factors separately, resulting in limited controllability or reduced visual quality. We revisit this problem from an image-first perspective, where high-quality human appearance is learned via image generation and used as a prior for video synthesis, decoupling appearance modeling from temporal consistency. We propose a pose- and viewpoint-controllable pipeline that combines a pretrained image backbone with SMPL-X-based motion guidance, together with a training-free temporal refinement stage based on a pretrained video diffusion model. Our method produces high-quality, temporally consistent videos under diverse poses and viewpoints. We also release a canonical human dataset and an auxiliary model for compositional human image synthesis. Code and data are publicly available at https://github.com/Taited/ReImagine.
title ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.19720