Animating the Uncaptured: Humanoid Mesh Animation with Video Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Millán, Marc Benedí San, Dai, Angela, Nießner, Matthias
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908276047740928
author Millán, Marc Benedí San
Dai, Angela
Nießner, Matthias
author_facet Millán, Marc Benedí San
Dai, Angela
Nießner, Matthias
contents Animation of humanoid characters is essential in various graphics applications, but requires significant time and cost to create realistic animations. We propose an approach to synthesize 4D animated sequences of input static 3D humanoid meshes, leveraging strong generalized motion priors from generative video models -- as such video models contain powerful motion information covering a wide variety of human motions. From an input static 3D humanoid mesh and a text prompt describing the desired animation, we synthesize a corresponding video conditioned on a rendered image of the 3D mesh. We then employ an underlying SMPL representation to animate the corresponding 3D mesh according to the video-generated motion, based on our motion optimization. This enables a cost-effective and accessible solution to enable the synthesis of diverse and realistic 4D animations.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15996
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Animating the Uncaptured: Humanoid Mesh Animation with Video Diffusion Models
Millán, Marc Benedí San
Dai, Angela
Nießner, Matthias
Graphics
Computer Vision and Pattern Recognition
Animation of humanoid characters is essential in various graphics applications, but requires significant time and cost to create realistic animations. We propose an approach to synthesize 4D animated sequences of input static 3D humanoid meshes, leveraging strong generalized motion priors from generative video models -- as such video models contain powerful motion information covering a wide variety of human motions. From an input static 3D humanoid mesh and a text prompt describing the desired animation, we synthesize a corresponding video conditioned on a rendered image of the 3D mesh. We then employ an underlying SMPL representation to animate the corresponding 3D mesh according to the video-generated motion, based on our motion optimization. This enables a cost-effective and accessible solution to enable the synthesis of diverse and realistic 4D animations.
title Animating the Uncaptured: Humanoid Mesh Animation with Video Diffusion Models
topic Graphics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.15996