Human Video Generation from a Single Image with 3D Pose and View Control

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Tiantian, Yao, Chun-Han, Hu, Tao, Reddy, Mallikarjun Byrasandra Ramalinga, Yang, Ming-Hsuan, Jampani, Varun
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917292834553856
author Wang, Tiantian
Yao, Chun-Han
Hu, Tao
Reddy, Mallikarjun Byrasandra Ramalinga
Yang, Ming-Hsuan
Jampani, Varun
author_facet Wang, Tiantian
Yao, Chun-Han
Hu, Tao
Reddy, Mallikarjun Byrasandra Ramalinga
Yang, Ming-Hsuan
Jampani, Varun
contents Recent diffusion methods have made significant progress in generating videos from single images due to their powerful visual generation capabilities. However, challenges persist in image-to-video synthesis, particularly in human video generation, where inferring view-consistent, motion-dependent clothing wrinkles from a single image remains a formidable problem. In this paper, we present Human Video Generation in 4D (HVG), a latent video diffusion model capable of generating high-quality, multi-view, spatiotemporally coherent human videos from a single image with 3D pose and view control. HVG achieves this through three key designs: (i) Articulated Pose Modulation, which captures the anatomical relationships of 3D joints via a novel dual-dimensional bone map and resolves self-occlusions across views by introducing 3D information; (ii) View and Temporal Alignment, which ensures multi-view consistency and alignment between a reference image and pose sequences for frame-to-frame stability; and (iii) Progressive Spatio-Temporal Sampling with temporal alignment to maintain smooth transitions in long multi-view animations. Extensive experiments on image-to-video tasks demonstrate that HVG outperforms existing methods in generating high-quality 4D human videos from diverse human images and pose inputs.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21188
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Human Video Generation from a Single Image with 3D Pose and View Control
Wang, Tiantian
Yao, Chun-Han
Hu, Tao
Reddy, Mallikarjun Byrasandra Ramalinga
Yang, Ming-Hsuan
Jampani, Varun
Computer Vision and Pattern Recognition
Recent diffusion methods have made significant progress in generating videos from single images due to their powerful visual generation capabilities. However, challenges persist in image-to-video synthesis, particularly in human video generation, where inferring view-consistent, motion-dependent clothing wrinkles from a single image remains a formidable problem. In this paper, we present Human Video Generation in 4D (HVG), a latent video diffusion model capable of generating high-quality, multi-view, spatiotemporally coherent human videos from a single image with 3D pose and view control. HVG achieves this through three key designs: (i) Articulated Pose Modulation, which captures the anatomical relationships of 3D joints via a novel dual-dimensional bone map and resolves self-occlusions across views by introducing 3D information; (ii) View and Temporal Alignment, which ensures multi-view consistency and alignment between a reference image and pose sequences for frame-to-frame stability; and (iii) Progressive Spatio-Temporal Sampling with temporal alignment to maintain smooth transitions in long multi-view animations. Extensive experiments on image-to-video tasks demonstrate that HVG outperforms existing methods in generating high-quality 4D human videos from diverse human images and pose inputs.
title Human Video Generation from a Single Image with 3D Pose and View Control
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.21188