Ponymation: Learning Articulated 3D Animal Motions from Unlabeled Online Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Keqiang, Litvak, Dor, Zhang, Yunzhi, Li, Hongsheng, Wu, Jiajun, Wu, Shangzhe
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914894475952128
author Sun, Keqiang
Litvak, Dor
Zhang, Yunzhi
Li, Hongsheng
Wu, Jiajun
Wu, Shangzhe
author_facet Sun, Keqiang
Litvak, Dor
Zhang, Yunzhi
Li, Hongsheng
Wu, Jiajun
Wu, Shangzhe
contents We introduce a new method for learning a generative model of articulated 3D animal motions from raw, unlabeled online videos. Unlike existing approaches for 3D motion synthesis, our model requires no pose annotations or parametric shape models for training; it learns purely from a collection of unlabeled web video clips, leveraging semantic correspondences distilled from self-supervised image features. At the core of our method is a video Photo-Geometric Auto-Encoding framework that decomposes each training video clip into a set of explicit geometric and photometric representations, including a rest-pose 3D shape, an articulated pose sequence, and texture, with the objective of re-rendering the input video via a differentiable renderer. This decomposition allows us to learn a generative model over the underlying articulated pose sequences akin to a Variational Auto-Encoding (VAE) formulation, but without requiring any external pose annotations. At inference time, we can generate new motion sequences by sampling from the learned motion VAE, and create plausible 4D animations of an animal automatically within seconds given a single input image.
format Preprint
id arxiv_https___arxiv_org_abs_2312_13604
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Ponymation: Learning Articulated 3D Animal Motions from Unlabeled Online Videos
Sun, Keqiang
Litvak, Dor
Zhang, Yunzhi
Li, Hongsheng
Wu, Jiajun
Wu, Shangzhe
Computer Vision and Pattern Recognition
We introduce a new method for learning a generative model of articulated 3D animal motions from raw, unlabeled online videos. Unlike existing approaches for 3D motion synthesis, our model requires no pose annotations or parametric shape models for training; it learns purely from a collection of unlabeled web video clips, leveraging semantic correspondences distilled from self-supervised image features. At the core of our method is a video Photo-Geometric Auto-Encoding framework that decomposes each training video clip into a set of explicit geometric and photometric representations, including a rest-pose 3D shape, an articulated pose sequence, and texture, with the objective of re-rendering the input video via a differentiable renderer. This decomposition allows us to learn a generative model over the underlying articulated pose sequences akin to a Variational Auto-Encoding (VAE) formulation, but without requiring any external pose annotations. At inference time, we can generate new motion sequences by sampling from the learned motion VAE, and create plausible 4D animations of an animal automatically within seconds given a single input image.
title Ponymation: Learning Articulated 3D Animal Motions from Unlabeled Online Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.13604